Repository navigation
Commit 72bb5fe
committed
perf: literal-prefix capture-extraction fast path
For anchored patterns of the shape
^<literal-prefix-set>([^X]+)X.*$ with replacement `${1}` (or `$1`)
capture 1's bounds are structurally trivial — skip the prefix, find the
terminator with memchr — so the engine doesn't need to track captures
at all.
Two changes work together:
1. A new `LiteralPrefixCapture` strategy in `regex-automata`'s meta
engine recognizes the shape via HIR walking (single-pattern only,
anchored at both ends, default flags, ASCII terminator, finite
literal-alternation prefix set capped at 32 variants). Strategy
methods extract the match and capture-1 slots directly with memchr,
bypassing PikeVM / BoundedBacktracker. Wires in alongside the
existing reverse strategies.
2. `Regex::replacen` gets a borrowed-output fast path for replacements
that are exactly `$N` / `${N}`. Detected via a new
`Replacer::single_capture_ref` method (default `None`, opted into
for `&str`/`String`/`Cow<str>`). For `limit == 1` with a match
covering the whole haystack, returns `Cow::Borrowed` of the
captured slice — no `Captures::expand`, no output string
allocation.
Bench (500k synthetic Referer rows, 5-iter mean, on the same machine):
Regex::replacen, q28 pattern, 80% match
before: 281 ms
after: 39 ms (7.3x)
Regex::replacen, ^key=([^,]+),.*$, 100% match
before: 113 ms
after: 27 ms (4.2x)
Tests: 257 / 257 pass (regex-automata --lib + --test integration, regex
--test integration). No regressions.1 parent 839d16b commit 72bb5fe
2 files changed
Lines changed: 537 additions & 2 deletions
0 commit comments