Skip to content

Use fancy instead of onig - #69

Open
Keats wants to merge 12 commits into
masterfrom
fancy-regex
Open

Keats wants to merge 12 commits into
masterfrom
fancy-regex

Conversation

@Keats

@Keats Keats commented Jun 25, 2026

Copy link
Copy Markdown
Contributor

No description provided.

@Keats

Keats commented Aug 29, 2026

Copy link
Copy Markdown
Contributor Author

Comparison with master using onig:

highlight jquery.js     time:   [269.70 ms 270.21 ms 270.77 ms]
                        change: [−2.6265% −2.3485% −2.1016%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
  4 (4.00%) high mild
  1 (1.00%) high severe

highlight simple.ts     time:   [10.640 ms 10.667 ms 10.694 ms]
                        change: [+90.973% +91.717% +92.496%] (p = 0.00 < 0.05)
                        Performance has regressed.
Found 10 outliers among 100 measurements (10.00%)
  4 (4.00%) low mild
  6 (6.00%) high mild

highlight multiple simple.ts
                        time:   [10.891 ms 10.951 ms 11.015 ms]
                        change: [+91.852% +93.049% +94.393%] (p = 0.00 < 0.05)
                        Performance has regressed.
Found 1 outliers among 100 measurements (1.00%)
  1 (1.00%) high mild

Benchmarking highlight rust.sample: Warming up for 3.0000 s
Warning: Unable to complete 100 samples in 5.0s. You may wish to increase target time to 7.2s, enable flat sampling, or reduce sample count to 50.
highlight rust.sample   time:   [1.4058 ms 1.4080 ms 1.4107 ms]
                        change: [+62.500% +62.972% +63.490%] (p = 0.00 < 0.05)
                        Performance has regressed.
Found 11 outliers among 100 measurements (11.00%)
  1 (1.00%) low severe
  1 (1.00%) low mild
  6 (6.00%) high mild
  3 (3.00%) high severe

highlight markdown.sample
                        time:   [20.404 ms 20.514 ms 20.634 ms]
                        change: [+24.745% +25.585% +26.389%] (p = 0.00 < 0.05)
                        Performance has regressed.
Found 8 outliers among 100 measurements (8.00%)
  7 (7.00%) high mild
  1 (1.00%) high severe

highlight javascript.sample
                        time:   [33.667 ms 33.862 ms 34.066 ms]
                        change: [−21.912% −21.130% −20.447%] (p = 0.00 < 0.05)
                        Performance has improved.

Some wins but mostly regression for cold highlight (fancy usually wins on warm highlight though).

@keith-hall I've tried letting Fable have another look at fancy-regex (relevant commits here: https://github.com/Keats/fancy-regex/commits/cold-perf/ they are all pretty short except 0de5d3e and ee914fe) and the results using that branch (minus ee914fe since it just got added) compared to the previous run above that uses current fancy-regex main branch:

highlight jquery.js     time:   [243.97 ms 244.51 ms 245.05 ms]
                        change: [−9.7641% −9.5109% −9.2558%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
  1 (1.00%) high mild

highlight simple.ts     time:   [4.5677 ms 4.5845 ms 4.6032 ms]
                        change: [−57.222% −57.021% −56.820%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 11 outliers among 100 measurements (11.00%)
  7 (7.00%) high mild
  4 (4.00%) high severe

highlight multiple simple.ts
                        time:   [4.6603 ms 4.6851 ms 4.7115 ms]
                        change: [−57.557% −57.218% −56.883%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 17 outliers among 100 measurements (17.00%)
  10 (10.00%) high mild
  7 (7.00%) high severe

highlight rust.sample   time:   [842.19 µs 849.80 µs 858.89 µs]
                        change: [−40.397% −39.966% −39.456%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 16 outliers among 100 measurements (16.00%)
  1 (1.00%) low mild
  9 (9.00%) high mild
  6 (6.00%) high severe

highlight markdown.sample
                        time:   [8.6155 ms 8.6910 ms 8.7875 ms]
                        change: [−58.062% −57.634% −57.092%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 13 outliers among 100 measurements (13.00%)
  1 (1.00%) high mild
  12 (12.00%) high severe

highlight javascript.sample
                        time:   [18.763 ms 18.909 ms 19.069 ms]
                        change: [−44.685% −44.160% −43.566%] (p = 0.00 < 0.05)
                        Performance has improved.

With those changes, fancy-regex beats oniguruma.
ee914fe is an additional 3-4% win, not sure if it's worth the complexity though.

@Keats

Keats commented Aug 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Ah and even with those patches, using RegexSet for short lived CLI like Zola is killing performance (around10x worse than the prefilter for the small highlight benchmarks). RegexSet wins over the prefilter approach only when already warm.

I'll add a flag to toggle for full RegexSet mode for something like a server running giallo but sadly I do have to keep the prefilter ;(

@keith-hall

Copy link
Copy Markdown

With those changes, fancy-regex beats oniguruma.

Yay, dream come true! Thanks, I will take a look at the changes and see what I can apply to fancy-regex.

even with those patches, using RegexSet for short lived CLI like Zola is killing performance (around10x worse than the prefilter for the small highlight benchmarks). RegexSet wins over the prefilter approach only when already warm

That is a shame... I wonder whether any strategies like JIT compilation (which I have seen other regex crates offering) would help and whether it is feasible...

@Keats

Keats commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

I'm wondering whether the prefilter should be in fancy-regex tbh.

I've added an option to use full regex set or prefilter in the last commit and you can play with it but you get a 15-60% perf improvement at the cost of 2-10x more memory usage in the full regex set usage. I don't think this tradeoff makes a lot of sense in practice, the improvement is rarely worth the memory usage for pretty much any fancy-regex user.

With a local Zola benchmark, I had a 100k pages site rendering in 10s with regexset and 15s with prefilter but the regexset used 12GB of memory while the prefilter used 6GB (and the difference would only grow if I highlight more languages).

@keith-hall

Copy link
Copy Markdown

Makes sense, feel free to prototype how it could look in fancy-regex and we'll see the impact and whether we want to include it. (Would it look identical to 7ed0d70, or perhaps make use of some internal methods, and/or perhaps be based on the seek pattern etc?)

@Keats

Keats commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

I'm not sure exactly where the limit would be. Should the prefilter from 7ed0d70 be moved (and probably optimized a bit, I used bool because it was simpler but we can do bit twiddling with u64 and it should be faster/smaller?) and then exposed and users can build their own walk or regexset? Or should fancy-regex expose a whole new approach like in matcher.rs of this PR?

For the prefilter: I would need to check what's available in the internal methods of fancy-regex, maybe we can simplify stuff.

If we wanted to move more to fancy-regex, we would need to add LazyRegex (only compile the re on demand the first time, not at instantiation time) and LazyRegexSet which is essentially what I'm doing there: try to do as much with LazyRegex and fall back to a RegexSet for the patterns when we can't prefilter. I think it makes sense to add but maybe I would be the only user in which case it probably wouldn't be worth it. Could Syntect benefit from it as well?

I can work on a PR, just let me know where the limit should be.

@keith-hall

keith-hall commented Sep 3, 2026 •

Copy link
Copy Markdown

Probably makes sense to do bit twiddling, to have the shared implementation as performant as possible...
Not sure, do we need to give that level of control? Maybe having as much in fancy-regex as possible makes sense so consumers don't need to reinvent the wheel. I think having lazy implementations could be beneficial for syntect too - likely once fancy-regex is officially faster than Oniguruma, there would be no reason to keep the Oniguruma code in syntect, especially as it is archived and unmaintained, thus making it easier to implement and justifying an overhaul.

@keith-hall

Copy link
Copy Markdown

We can put it behind a feature flag in case others don't need it

@Keats

Keats commented Sep 6, 2026

Copy link
Copy Markdown
Contributor Author

In one commit: Keats/fancy-regex@dfdb09c
Tests and docs missing, I'll add them after we decide what we want to keep/remove.

@keith-hall

Copy link
Copy Markdown

As it will be behind a feature likely only Giallo will use, I'd be tempted to say include what you need, whatever makes sense to have in fancy-regex...

@Keats

Keats commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

If it's only giallo then I will just keep it in giallo and if someone else needs it, the branch is always there to add it in fancy-regex

@Keats
Keats force-pushed the fancy-regex branch 2 times, most recently from 8a2c312 to 2bd5f00 Compare September 14, 2026 12:01
@Keats

Keats commented Sep 14, 2026

Copy link
Copy Markdown
Contributor Author

FYI @keith-hall the ClassSeq changes from https://github.com/Keats/fancy-regex/commits/cold-perf/ give the cold benchmarks of giallo a good range of 0 to 30% speedup. It also does improve a bit some warm benchmarks, some -6/8% for JS/TS.

@keith-hall

Copy link
Copy Markdown

I haven't forgotten, its on my todo list - I feel those changes are complicated and I will need some time to fully understand them before I merge them in. Thanks for the reminder and your patience 🙂 (I saw that the rust coreutils implementation wanted to switch to fancy-regex, so I ended up spending some time to add leftmost-longest semantics as an option, which took some of my time 😉)

@Keats

Keats commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

To make it simple, I've put all the unapplied fancy-regex patches (including the ClassSeq ones) into https://github.com/Keats/fancy-regex/tree/class-seq
It also contains 2 more commits that are pretty small and pure wins it seems.
Edit: and another one
Edit: and another one

Use changes from fancy-regex lazy branch

Prebuild bytesets and update benches

Do not build a regexset for a single regex
@Keats

Keats commented Sep 20, 2026

Copy link
Copy Markdown
Contributor Author

I wanted to help you by sending individual PRs but I don't want to send slop and I don't know the fancy-regex crate enough to write it myself so I'm not sure I'm actually helping. I'll list individual branches with AI generated commits with benches/tests for the things that move the needle a lot rather than one branch so it's easier to review:

@keith-hall

Copy link
Copy Markdown

Thanks, I applied the lookbehind one at fancy-regex/fancy-regex#285

@keith-hall

Copy link
Copy Markdown

And I opened a PR for the ClassSeq changes at fancy-regex/fancy-regex#290, will force me to review it ;)

@Keats

Keats commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

@keith-hall benchmarks comparing master with onig and this branch with fancy-regex/class-seq2: https://gist.github.com/Keats/85dfd018d732fdf77d86a3a4d2ff9606

This branch has a lot of optimizations that could certainly be used with onig as well so not an perfect comparison but good enough.

The 2 biggest regressions are PHP (rust-lang/regex#1396 for the upstream fix) and shell (slevithan/oniguruma-parser#28 the shellscript has a pretty bad pattern for fancy-regex that is not optimized). I'll hardcode replacements for those in giallo until it gets fixed upstream but it's almost a clear win everywhere. The jquery file highlight is even 2x faster than syntect on my machine, that's almost entirely the required bytes filter at work.

Look at those warm numbers 👀

@keith-hall

Copy link
Copy Markdown

Nice one, thanks for sharing. I wonder whether the required byte filter would help improve performance for findutils: uutils/findutils#864

Do you think it is worth trying to apply a similar suffix extraction optimization in fancy-regex?

How is the memory usage now btw?

@Keats

Keats commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

I wonder whether the required byte filter would help improve performance for findutils: uutils/findutils#864

Worth a try. It depends a lot on the patterns, eg it was a huge win for JS but almost a no-op on a few other lang. Do they use a single regex? Multiple? I haven't looked at their code. They also should check if they want to disable prefilters etc

Do you think it is worth trying to apply a similar suffix extraction optimization in fancy-regex?

Maybe? Shiki runs https://github.com/slevithan/oniguruma-parser/tree/main/src/optimizer on all the textmate grammars so giallo would not benefit from it and I don't know who else is using crazy regexes like textmate/sublime.

How is the memory usage now btw?

Not too bad. Something like 50% higher than onig still. I've removed the public RegexSet strategy though as with all the optimizations it was same speed or slower than the walk while using twice the memory. I still use them internally in a few places though.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants