Repository navigation
Conversation
A `grep-captures` benchmark modeling ClickBench Q28, which extracts the
URL host out of HTTP `Referer` values via
REGEXP_REPLACE("Referer", '^https?://(?:www\.)?([^/]+)/.*$', '\1')
The regex is anchored at both ends, the prefix reduces to a finite set
of literal byte alternatives, the capture is a single `[^/]+`, and the
tail is `.*$`. Patterns of this shape are common for "extract a field
bounded by a known delimiter" tasks — host extraction here, but also
log fields, single-line CSV-ish parsing, path components, and similar.
Haystack: 20,000 newline-separated synthetic Referer URLs (~800 KiB),
deterministically generated with ~80% matching the pattern. Expected
capture count: 32,000 (16,000 matching lines × {group 0, group 1}).
Also adds a small `compile` benchmark for the same regex.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
08b232a to
76ea776
Compare
|
This is amazing work. The write-up here is especially lovely. Thank you for that. I see your PR on With that said... I don't have a good sense of how well the strategy infrastructure scales. This isn't totally fair to pin on you, but could you run the Thank you. :-) <3 |
I must admit this was mostly done by Claude (with some bit of steering / cleaning up of course) and I am not familiar with the code base so feel free to let it "cool down" a bit ;) or to let it inspire another implementation The compilation time is within noise (earlier comment said it was slower (~16 vs ~70 us), but that was compared to rust/regex 1.7.3 which seems to have some lower overhead). 1.12.2 has similar overhead: |
A new
grep-capturesbenchmark modeling ClickBench Q28 (analytic SQL benchmark), which extracts the URL host from HTTPReferervalues:Currently
rust/regexis about 6.5x slower thanjavascript/v8on this benchmark.results