Skip to content

Skipping whitespace tokens #156

Description

@jarble

Is it possible to skip tokens when defining a lexer?
I want to split a string into a list of tokens without whitespace, but I don't know if Moo can do this:

Input string:

  • "while ( a < 3 ) { a += 1; }"

List of tokens:

  • ["while","(","a","<","3",")","{","a","+=","1",",";","}"]

Activity

nathan commented on May 18, 2021

@nathan
Collaborator
const moo = require('moo')
const lex = moo.compile({
  ws: {match: /\p{White_Space}+/u, lineBreaks: true},
  word: /\p{XID_Start}\p{XID_Continue}*/u,
  op: moo.fallback,
})
;[...lex.reset('while ( a < 3 ) { a += 1; }')]
.filter(t => t.type !== 'ws')
.map(t => t.value)

jarble commented on May 18, 2021

@jarble
Author

@nathan The documentation doesn't describe this feature: does it need to be updated?

tjvr commented on May 18, 2021

@tjvr
Collaborator

The documentation needs to be updated to document moo.fallback (see #112).

As for the rest, I think Nathan's just demonstrating that since a moo lexer object is an Iterator, you can use filter() and map() which are built-in to JavaScript.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions