Repository navigation
Conversation
For anchored patterns of the shape
^<literal-prefix-set>([^X]+)X.*$ with replacement `${1}` (or `$1`)
capture 1's bounds are structurally trivial — skip the prefix, find the
terminator with memchr — so the engine doesn't need to track captures
at all.
Two changes work together:
1. A new `LiteralPrefixCapture` strategy in `regex-automata`'s meta
engine recognizes the shape via HIR walking (single-pattern only,
anchored at both ends, default flags, ASCII terminator, finite
literal-alternation prefix set capped at 32 variants). Strategy
methods extract the match and capture-1 slots directly with memchr,
bypassing PikeVM / BoundedBacktracker. Wires in alongside the
existing reverse strategies.
2. `Regex::replacen` gets a borrowed-output fast path for replacements
that are exactly `$N` / `${N}`. Detected via a new
`Replacer::single_capture_ref` method (default `None`, opted into
for `&str`/`String`/`Cow<str>`). For `limit == 1` with a match
covering the whole haystack, returns `Cow::Borrowed` of the
captured slice — no `Captures::expand`, no output string
allocation.
Bench (500k synthetic Referer rows, 5-iter mean, on the same machine):
Regex::replacen, q28 pattern, 80% match
before: 281 ms
after: 39 ms (7.3x)
Regex::replacen, ^key=([^,]+),.*$, 100% match
before: 113 ms
after: 27 ms (4.2x)
Tests: 257 / 257 pass (regex-automata --lib + --test integration, regex
--test integration). No regressions.
The fast path validated the *whole* haystack as UTF-8 whenever either the capture class or the `.*` class was a Unicode class. That is too strict when UTF-8 mode is disabled (e.g. `meta::Builder` with `syntax::Config::new().utf8(false)`): the literal prefix and any `(?-u:...)` span may legitimately contain invalid UTF-8. For example, `^((?-u:[^/])+)/.*$` failed to match `b"\xA9/"` and `^(?-u:\xFF)([^/]+)/.*$` failed to match `b"\xFFab/x"`. Now we record separately whether the capture and the tail need UTF-8, validate exactly those spans, and treat a failure as "this prefix does not match" instead of "no match", since another prefix yields a different capture span. Because the validation is now exact, Unicode classes no longer need to be gated on `utf8_empty`, which lets `regex::bytes::Regex` (Unicode mode) use the fast path too. Found by differential fuzzing against the PikeVM. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKuwE1LGom89mpJoo8YDtj
The memchr2 loop over the terminator and `\n` stopped at each newline only to resume the scan, because `[^X]+` may contain newlines. The result is always the first terminator, so a plain memchr is equivalent and simpler. (Benchmarks are neutral on URL-shaped inputs.) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKuwE1LGom89mpJoo8YDtj
For `replace` (limit == 1): * With a `$N` replacement for a small N, search with stack-allocated slots via `meta::Regex::search_slots` instead of allocating a `Captures` value. * With any other template, call `captures` directly instead of going through `captures_iter`, which allocates a `Captures` on construction and then clones it for each match. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKuwE1LGom89mpJoo8YDtj
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This branch is rust-lang#1350 (6 commits) with 3 follow-up commits on top, for cherry-picking into the upstream PR. Only the last 3 commits are new.
Follow-up commits
automata: validate only Unicode-class spans in literal prefix capture(correctness)The fast path validated the whole haystack as UTF-8 whenever the capture class or the
.*class was a Unicode class. When UTF-8 mode is off, the literal prefix and any(?-u:…)span may legitimately contain invalid UTF-8, so valid matches were dropped:^((?-u:[^/])+)/.*$onb"\xA9/"returned no match; PikeVM matches0..2^([^\n]+)\n(?-u:.)*$onb"x\na\xA9bb"returned no match^(?-u:\xFF)([^/]+)/.*$onb"\xFFab/x"returned no matchThis only happens through
regex_automata::metawithsyntax.utf8(false)andutf8_empty(true). Theregexcrate's defaults never reach it. The fix validates exactly the capture span and the tail span. A failure now skips to the next prefix instead of returningNone. Because validation is now exact, Unicode classes no longer need gating onutf8_empty, soregex::bytes::Regexalso gets the fast path (Q28 bytes replace: 1411 → 298 ns).automata: use a single memchr to find the capture terminator(simplification)[^X]+may contain\n, so thememchr2(X, '\n')loop always ends at the firstX. A plainmemchrgives the same result. Performance is unchanged.regex: avoid Captures allocation in single-match replace(perf)For
replace(limit == 1): a$Nreplacement with N < 4 uses stack slots viameta::Regex::search_slots. Other templates callcaptures()directly instead ofcaptures_iter, which allocates aCapturesand then clones it per match.Benchmarks
The harness is a standalone binary, not committed. It uses 200k synthetic Referer-style URLs (60% https, 50%
www., 20–160 bytes, some non-ASCII), pinned to one core, and reports the best of 3×7 runs.replace(s, "$1")replace(s, "x${1}")capturesis_matchbytes::Regexreplacebytes::Regex(?-u)replace(\w+)@first-matchreplace "$1"(not Q28 shape)replace_all "$1", 40k matchesAblations on rust-lang#1350 itself:
Input::new_utf8, so the fast path always validates: +30 ns on short URLs, 2x slower on 2 KB URLs. The new public API is worth something.single_capture_ref: Q28 replace goes 164 → 256 ns.Testing
tests/misc.rs,tests/replace.rs).utf8_emptyon/off) comparingmeta::Regexcapture spans againstPikeVM. Before fix: 4957 mismatches in 2.8M checks. After: 0 in about 17M checks across 4 seeds.cargo testpasses forregexand forregex-automata --all-features.--no-default-featuresbuilds pass.Other review notes on rust-lang#1350 (no code change here)
Input::new_utf8(regex-automata) andReplacer::single_capture_ref(regex) are both new public items. Upstream may prefer#[doc(hidden)]or a narrower design.MAX_PREFIX_VARIANTScomment says 32Box<[u8]>fit "on one cache line". They take 512 bytes.(?:H|h)ttps?becomes[Hh]ttps?), so they are not recognized. Expanding small byte classes inprefix_variantswould widen coverage.Coreis still built, but it is only used forwhich_overlapping_matchesand the cache. Build time and memory are unchanged, which is fine.🤖 Generated with Claude Code
https://claude.ai/code/session_01JKuwE1LGom89mpJoo8YDtj
Generated by Claude Code