Skip to content

Commit 72bb5fe

Browse files
committed
perf: literal-prefix capture-extraction fast path
For anchored patterns of the shape ^<literal-prefix-set>([^X]+)X.*$ with replacement `${1}` (or `$1`) capture 1's bounds are structurally trivial — skip the prefix, find the terminator with memchr — so the engine doesn't need to track captures at all. Two changes work together: 1. A new `LiteralPrefixCapture` strategy in `regex-automata`'s meta engine recognizes the shape via HIR walking (single-pattern only, anchored at both ends, default flags, ASCII terminator, finite literal-alternation prefix set capped at 32 variants). Strategy methods extract the match and capture-1 slots directly with memchr, bypassing PikeVM / BoundedBacktracker. Wires in alongside the existing reverse strategies. 2. `Regex::replacen` gets a borrowed-output fast path for replacements that are exactly `$N` / `${N}`. Detected via a new `Replacer::single_capture_ref` method (default `None`, opted into for `&str`/`String`/`Cow<str>`). For `limit == 1` with a match covering the whole haystack, returns `Cow::Borrowed` of the captured slice — no `Captures::expand`, no output string allocation. Bench (500k synthetic Referer rows, 5-iter mean, on the same machine): Regex::replacen, q28 pattern, 80% match before: 281 ms after: 39 ms (7.3x) Regex::replacen, ^key=([^,]+),.*$, 100% match before: 113 ms after: 27 ms (4.2x) Tests: 257 / 257 pass (regex-automata --lib + --test integration, regex --test integration). No regressions.
1 parent 839d16b commit 72bb5fe

2 files changed

Lines changed: 537 additions & 2 deletions

File tree

0 commit comments

Comments
 (0)