Repository navigation
Conversation
IgnoreCase matching resolved equivalent characters with unicode.SimpleFold, but .NET groups characters by their invariant lowercase form. The two agree across ASCII and diverge outside it, so patterns matched characters .NET considers unrelated (and vice versa): "s" matched U+017F LATIN SMALL LETTER LONG S "σ" matched U+03C2 GREEK SMALL LETTER FINAL SIGMA "[a-z-[aeiou]]" matched 'A', because subtracted sets were not folded "(.)\1" matched "iİ", because backrefs compared with unicode.ToLower tryFindCaseEquivalences now looks the character up in a table built once from unicode.CaseRanges and keyed by invariant lowercase, addCaseEquivalences recurses into the subtracted set, and the parser, the runner's literal and backreference comparisons, and addLowercase go through InvariantToLower, which differs from unicode.ToLower only for U+0130. The lcTable row mapping U+0130 to 'i' was a culture-sensitive mapping the invariant culture does not make, so it is gone. Checked differentially against .NET 8 with CultureInvariant over 34,053 generated pattern/input cases: 85 divergences before, 0 after. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks for this info! Looks like this behavior is split amongst various regex engines (tested only on regex101):
Is there a case to be made that one behavior is more correct than the other? Did this come up in real usage? |
|
This came out of differential testing rather than a bug report. I ran a corpus of 34,053 (pattern, input) pairs through regexp2 and through a real .NET 8 runtime and diffed the results; 85 disagreed, all under IgnoreCase and all outside ASCII. I can share the harness if it's useful. It's not really a question of Unicode correctness, since both relations are defensible, but of what this library promises. Regexp2 is positioned as a port of .NET's regex engine, and people use it over the standard library to run patterns written for .NET and get the same matches. Currently, under IgnoreCase they don't, and the discrepancies are the type that can be unnoticed until they matter: If you'd prefer to keep SimpleFold for the RE2 compatibility mode, since that mode's users expect Go semantics, I'm happy to gate the equivalence table on the default .NET mode only and leave RE2 mode as it is. Let me know which you'd rather have and I'll update the PR. |
|
The unicode and string matching behavior has never exactly matched .NET because of the differences in UTF8 strings in Go vs .NET (as you pointed out in your However since this is difference is so niche I'm not against making this compatibility move, but a few things I'd like to see:
|
The divergence
IgnoreCaseresolves equivalent characters withunicode.SimpleFold. .NET does not useUnicode simple case folding — it groups characters into equivalence classes by their
invariant lowercase form. The two relations agree across ASCII and diverge outside it,
so regexp2 currently matches characters .NET considers unrelated, and misses some it
considers equivalent.
Concrete cases where regexp2 and .NET disagree today (all with
RegexOptions.IgnoreCase | RegexOptions.CultureInvarianton the .NET side):sſ(U+017F LATIN SMALL LETTER LONG S)σς(U+03C2 GREEK SMALL LETTER FINAL SIGMA)[a-z-[aeiou]]A(.)\1iİ(U+0069, U+0130)U+017F lowercases to itself, so it shares no lowercase form with
sorS; simple foldingrelates all three. Same story for U+03C2 against U+03C3/U+03A3. The subtraction case is a
separate bug: the subtracted set was never case-folded, so
[a-z-[aeiou]]foldeda-zupto
A-Zbut kept subtracting only the lowercase vowels. And backreference comparison usedunicode.ToLower, which folds U+0130 ontoi— a culture-sensitive mapping the invariantculture does not make.
How .NET defines it
dotnet/runtime'sRegexCaseEquivalencesholds the equivalence table the regex engine consults, generated from the case mappings of
the invariant culture.
RegexCharClassbuilds its sets from it. The upshot is that twocharacters are case-equivalent exactly when they have the same invariant lowercase form —
and that
ToLowerInvariantleaves U+0130 (LATIN CAPITAL LETTER I WITH DOT ABOVE) alone,where
unicode.ToLowermaps it toi.The fix
syntax/charclass.gotryFindCaseEquivalencesnow looks the character up in a table keyed by invariantlowercase, built once lazily from
unicode.CaseRanges(1448 groups, 2880 runes), ratherthan walking
unicode.SimpleFold.addCaseEquivalencesrecurses intoc.sub, so a subtracted set is folded along with theset it narrows.
InvariantToLower, which isunicode.ToLowerexcept that it leaves U+0130 alone.addLowercasegoes through it, so the trailing lowercasing pass inscanCharSetnolonger collapses U+0130 before equivalences are computed.
{'İ', 'İ', LowercaseSet, 0x0069}row fromlcTablefor the samereason.
syntax/parser.go— single-character literals under(?i)lowercase withInvariantToLower.runner.go—runematchandrefmatchcompare withInvariantToLower, so literal runsand backreferences use the same relation as the compiled sets.
The Boyer-Moore prefix in
syntax/prefix.goand thehelpersindex-of routines still useunicode.ToLower. That relation is strictly coarser than the new one, so they stay soundas prefilters — their extra candidates are rejected by the engine. I left them alone to
keep the diff focused; happy to tighten them in a follow-up if you'd prefer.
How it was verified
I ran regexp2 differentially against a .NET 8 console oracle (
RegexwithRegexOptions.IgnoreCase | RegexOptions.CultureInvariant) over 34,053 generatedpattern/input cases — single literals, ranges, class subtraction, unicode categories and
backreferences across the BMP. 85 cases diverged before this change and 0 after. Happy
to share the harness if it's useful.
CultureInvariantmatters here: without it the runtime applies the ambient culture's "Ibehaviour" and reports U+0130 as equivalent to
i, which produced one wrong expectationbefore the oracle was corrected.
Supplementary-plane characters are deliberately out of scope. .NET is UTF-16, so
[\U00010400]is a two-unit sequence there and[\U00010400-\U00010402]is rejected as areversed range; regexp2 is rune-based and behaves better than .NET in that region already.
Tests
regexp_ignorecase_test.goadds 30 table-driven cases in six groups — equivalence classes,equivalences retained, class subtraction, category subtraction, ranges, backreferences —
each commented with the .NET behaviour it pins. All 30 fail on master and pass with the
change.
gofmt -l .andgo vet ./...are clean, andgo test -count=1 ./...passes across allfour packages including the PCRE/RE2/Rust corpus tests.
🤖 Generated with Claude Code