Skip to content

Match .NET's case-equivalence rules under IgnoreCase - #111

Open
sawy3r wants to merge 1 commit into
dlclark:masterfrom
sawy3r:ignorecase-dotnet-equivalence
Open

sawy3r wants to merge 1 commit into
dlclark:masterfrom
sawy3r:ignorecase-dotnet-equivalence

Conversation

@sawy3r

@sawy3r sawy3r commented Sep 17, 2026

Copy link
Copy Markdown

The divergence

IgnoreCase resolves equivalent characters with unicode.SimpleFold. .NET does not use
Unicode simple case folding — it groups characters into equivalence classes by their
invariant lowercase form. The two relations agree across ASCII and diverge outside it,
so regexp2 currently matches characters .NET considers unrelated, and misses some it
considers equivalent.

Concrete cases where regexp2 and .NET disagree today (all with RegexOptions.IgnoreCase | RegexOptions.CultureInvariant on the .NET side):

pattern input regexp2 .NET
s ſ (U+017F LATIN SMALL LETTER LONG S) match no match
σ ς (U+03C2 GREEK SMALL LETTER FINAL SIGMA) match no match
[a-z-[aeiou]] A match no match
(.)\1 iİ (U+0069, U+0130) match no match

U+017F lowercases to itself, so it shares no lowercase form with s or S; simple folding
relates all three. Same story for U+03C2 against U+03C3/U+03A3. The subtraction case is a
separate bug: the subtracted set was never case-folded, so [a-z-[aeiou]] folded a-z up
to A-Z but kept subtracting only the lowercase vowels. And backreference comparison used
unicode.ToLower, which folds U+0130 onto i — a culture-sensitive mapping the invariant
culture does not make.

How .NET defines it

dotnet/runtime's
RegexCaseEquivalences
holds the equivalence table the regex engine consults, generated from the case mappings of
the invariant culture. RegexCharClass builds its sets from it. The upshot is that two
characters are case-equivalent exactly when they have the same invariant lowercase form —
and that ToLowerInvariant leaves U+0130 (LATIN CAPITAL LETTER I WITH DOT ABOVE) alone,
where unicode.ToLower maps it to i.

The fix

syntax/charclass.go

  • tryFindCaseEquivalences now looks the character up in a table keyed by invariant
    lowercase, built once lazily from unicode.CaseRanges (1448 groups, 2880 runes), rather
    than walking unicode.SimpleFold.
  • addCaseEquivalences recurses into c.sub, so a subtracted set is folded along with the
    set it narrows.
  • New InvariantToLower, which is unicode.ToLower except that it leaves U+0130 alone.
  • addLowercase goes through it, so the trailing lowercasing pass in scanCharSet no
    longer collapses U+0130 before equivalences are computed.
  • Dropped the {'İ', 'İ', LowercaseSet, 0x0069} row from lcTable for the same
    reason.

syntax/parser.go — single-character literals under (?i) lowercase with
InvariantToLower.

runner.go — runematch and refmatch compare with InvariantToLower, so literal runs
and backreferences use the same relation as the compiled sets.

The Boyer-Moore prefix in syntax/prefix.go and the helpers index-of routines still use
unicode.ToLower. That relation is strictly coarser than the new one, so they stay sound
as prefilters — their extra candidates are rejected by the engine. I left them alone to
keep the diff focused; happy to tighten them in a follow-up if you'd prefer.

How it was verified

I ran regexp2 differentially against a .NET 8 console oracle (Regex with
RegexOptions.IgnoreCase | RegexOptions.CultureInvariant) over 34,053 generated
pattern/input cases — single literals, ranges, class subtraction, unicode categories and
backreferences across the BMP. 85 cases diverged before this change and 0 after. Happy
to share the harness if it's useful.

CultureInvariant matters here: without it the runtime applies the ambient culture's "I
behaviour" and reports U+0130 as equivalent to i, which produced one wrong expectation
before the oracle was corrected.

Supplementary-plane characters are deliberately out of scope. .NET is UTF-16, so
[\U00010400] is a two-unit sequence there and [\U00010400-\U00010402] is rejected as a
reversed range; regexp2 is rune-based and behaves better than .NET in that region already.

Tests

regexp_ignorecase_test.go adds 30 table-driven cases in six groups — equivalence classes,
equivalences retained, class subtraction, category subtraction, ranges, backreferences —
each commented with the .NET behaviour it pins. All 30 fail on master and pass with the
change.

gofmt -l . and go vet ./... are clean, and go test -count=1 ./... passes across all
four packages including the PCRE/RE2/Rust corpus tests.

🤖 Generated with Claude Code

IgnoreCase matching resolved equivalent characters with unicode.SimpleFold,
but .NET groups characters by their invariant lowercase form. The two agree
across ASCII and diverge outside it, so patterns matched characters .NET
considers unrelated (and vice versa):

  "s" matched U+017F LATIN SMALL LETTER LONG S
  "σ" matched U+03C2 GREEK SMALL LETTER FINAL SIGMA
  "[a-z-[aeiou]]" matched 'A', because subtracted sets were not folded
  "(.)\1" matched "iİ", because backrefs compared with unicode.ToLower

tryFindCaseEquivalences now looks the character up in a table built once from
unicode.CaseRanges and keyed by invariant lowercase, addCaseEquivalences
recurses into the subtracted set, and the parser, the runner's literal and
backreference comparisons, and addLowercase go through InvariantToLower, which
differs from unicode.ToLower only for U+0130. The lcTable row mapping U+0130
to 'i' was a culture-sensitive mapping the invariant culture does not make, so
it is gone.

Checked differentially against .NET 8 with CultureInvariant over 34,053
generated pattern/input cases: 85 divergences before, 0 after.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@dlclark

dlclark commented Sep 17, 2026

Copy link
Copy Markdown
Owner

Thanks for this info! Looks like this behavior is split amongst various regex engines (tested only on regex101):

  • Go, Rust, Python, grep -- SimpleFold
  • PCRE2, Java, C#, Javascript -- not SimpleFold

Is there a case to be made that one behavior is more correct than the other? Did this come up in real usage?

@sawy3r

sawy3r commented Sep 17, 2026

Copy link
Copy Markdown
Author

This came out of differential testing rather than a bug report. I ran a corpus of 34,053 (pattern, input) pairs through regexp2 and through a real .NET 8 runtime and diffed the results; 85 disagreed, all under IgnoreCase and all outside ASCII. I can share the harness if it's useful.

It's not really a question of Unicode correctness, since both relations are defensible, but of what this library promises. Regexp2 is positioned as a port of .NET's regex engine, and people use it over the standard library to run patterns written for .NET and get the same matches. Currently, under IgnoreCase they don't, and the discrepancies are the type that can be unnoticed until they matter: ^[a-z]+$ with IgnoreCase accepts ſ (long s) and ς (final sigma) here but rejects them in .NET, and İ (U+0130) fails to match itself in one direction. For anyone using a ported pattern as an allow-list that's a silent behaviour change. .NET itself settled on the table-driven equivalence approach in .NET 7 (RegexCaseEquivalences in dotnet/runtime), so it's the current behaviour rather than a legacy quirk.

If you'd prefer to keep SimpleFold for the RE2 compatibility mode, since that mode's users expect Go semantics, I'm happy to gate the equivalence table on the default .NET mode only and leave RE2 mode as it is. Let me know which you'd rather have and I'll update the PR.

@dlclark

dlclark commented Sep 17, 2026

Copy link
Copy Markdown
Owner

The unicode and string matching behavior has never exactly matched .NET because of the differences in UTF8 strings in Go vs .NET (as you pointed out in your [\U00010400-\U00010402] example). Since the library has its own users now I do have to consider backward compatibility as a factor, not just .NET compatibility.

However since this is difference is so niche I'm not against making this compatibility move, but a few things I'd like to see:

  • RE2 mode should maintain existing behavior
  • Documentation update to make this change clear (as an RE2 mode difference)
  • Existing Benchmark suite run between old and new code to confirm minimal performance impact

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants