Skip to content

About

Independent determinacy audit of SWE-rebench 2026_03: 14.5% pointer-checkable claimable spine (7 airtight + 9 codebase-plural). Produced with the determinacy tool.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

23 Commits

Folders and files

Repository files navigation

swe-rebench-audit

An independent determinacy audit of SWE-rebench (Nebius, arXiv:2505.20411), the May 2026 leaderboard batch: the 2026_03 split on HuggingFace (n=110, PRs created March through mid-May 2026). The audited set is that split verbatim, with identical instance_ids, so there is no ambiguity about which tasks this covers. For each task it asks whether the issue text determines the behavior the hidden test grades, and reports the subset where it does not, with receipts you can re-check.

This is a measurement meant to inform evaluation design, scoped to one batch under one method. It is not a claim that the benchmark is bad. (Disclosure: we run SWE-rebench's official harness on our own science track.)

Produced with determinacy. Reproduce with determinacy run examples/swebench-rebench.toml.

Result

The audit separates two epistemic tiers, and they should be read separately.

Hard floor: 7 / 110 airtight. The hidden test grades a specific constant (a class name, a key, an exact message) that is absent from the issue text and from the repository at base_commit, present only in the gold patch and the test. These are mechanically checkable once the witness is identified: clone the repo, rg the constant, get no matches. For example, tobymao__sqlglot-7479 asserts isinstance(..., exp.AIEmbed), and the class name AIEmbed appears nowhere a solver could read it. Each case ships a single re-runnable command in its RECEIPT.md.

Additional rebuttable tier: +9 / 110 codebase-plural. The repository itself makes the relevant choice in ≥2 conflicting live ways while the issue is silent, so a solver reading the codebase could land on either. Each cites ≥2 grep-verified precedents and survives a comparability screen (an independent pass that excludes superficially-similar precedents). This tier still rests on a model judgment that the precedents are genuinely comparable, so read it as separable and rebuttable.

Combined: 16 / 110, of which 9 depend on comparability judgments. For context, the screens that find candidates flag much more (about 29% of graded behaviors are prose-silent at the behavior level; about 33% of tasks trip a two-rater prose-only screen), but screens over-flag and are reported only as upper bounds.

tier count evidence contestable?
airtight 7/110 constant absent from prose and codebase (one grep) no
codebase-plural 9/110 ≥2 conflicting live precedents, comparability-screened yes (model judgment)
combined 16/110 partially

Rebuttal boundary: disagreement with any plural case does not affect the 7/110 airtight floor. The floor is the stable result; the plural tier is an additional, separable estimate.

What it suggests, and what you can do with it

The airtight floor isolates concrete tasks where the hidden test grades a value the issue never stated. Even if you reject every plural case, those 7 are specific instances that could be addressed by clarifying the issue prose, adjusting the hidden test, or setting a policy for such cases. A practical first pass would be to review the 7 airtight receipts and decide case-by-case which of those to do, or to mark the task as intentionally codebase-resolving. More broadly, the indeterminate fraction measured here is the kind of signal under which the determinacy-weighted reporting explored for SWE-bench Pro (score weighted by determinacy; publish the underdetermined set) becomes worth considering for SWE-rebench. This audit measures that fraction; whether to act on it is the maintainers' call.

Honesty

  • Screens over-flag, disclosed as upper bounds. A test that parametrizes one decision across many inputs inflates the behavior-level count, so only the airtight and comparability-screened tiers are reported as findings.
  • The airtight subset is the stable piece. It is deterministic to check (a grep) and independent of any model. The plural tier and the screens involve model steps and should be read with that in mind.
  • A single pass under-counts. The method proposes one witness per case, so any one run surfaces only some airtight cases; reruns are designed to increase recall, not precision. ~29 codebase hypotheses went unprocessed (clone or agent failures) and could add to the certified set if processed.
  • Evaluation-design claim, not maintainer-intent. The audit shows the presented prose underdetermines the graded answer under blind evaluation (the agent-facing condition). It does not claim maintainers lacked intent.

Receipts and reproduction

Every task has a per-problem receipt: the verdict, the evidence, and for airtight cases a copy-paste git clone + rg that returns no matches. Airtight witnesses: data/cases/*/AMBIGUITY_WITNESS.md. Per-behavior coverage tables: data/attribution/. Reproduce end-to-end with determinacy. Method depth and the decision trail: PREREGISTRATION.md, WORKLOG.md, notes/. Full methodology reference: swebench-pro-audit.

About

Independent determinacy audit of SWE-rebench 2026_03: 14.5% pointer-checkable claimable spine (7 airtight + 9 codebase-plural). Produced with the determinacy tool.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages