Most benchmark wins don't survive a second look.
I find the ones that don't — before you publish them. Independent, reproducible audits of the numbers your model, your eval, or your leaderboard is about to stake its credibility on. If a claim is real, you get a clean bill of health you can cite. If it isn't, you find out from me, quietly, not from a reviewer or a competitor on launch day.
Eleven audits. Ten numbers that overstated — one that held.
Each was recomputed from public, released data. Ten target the one summary sentence that overpromised — a judge's bias, a lucky subset, a mechanical "law", an order effect, a base rate, a hidden second dimension, a crown won by luck. The eleventh is the control: a #1 I tried to break and couldn't, so I certify it. An auditor who only ever says no isn't checking anything. This is exactly the work I'd do on yours.
The judge preferred longer answers and its own family.
A widely-used LLM-as-judge, audited on its own released pairwise verdicts.
judge's own model family wins 71.5%
both p ≈ 0, fully reproducible from released data
0.978 AUC on the benchmark. About 12% precision in the clinic.
The same model, reprojected from a balanced test set to the prevalence an operator actually faces.
at 1% prevalence ≈12% precision — 7 of 8 alarms false
at 0.1% ≈1%, deterministic from Bayes
The ranking is one column. The data has up to six dimensions.
Parallel analysis on the model × item matrix, testing whether "better" is really a direction.
ProteinGym k ≈ 3 — second axis has a nameable mechanism
SWE-bench k ≈ 2, each gated by a permutation null
On five of seven boards, the announced #1 is more likely wrong than right.
Not a defect anyone can fix — sorting on score picks the largest error along with the largest ability.
winning score inflated +0.97 to +2.66 points
control 0.0% on the one board separated by 8 SE
Answering "D" to everything outranks 22 of the 29 published models.
The answer key is not uniform. A constant answer reads nothing and still lands above the board median.
below chance 11 of 29 published runs
bound same key on a later board beats only 13%
423 published records carry a standard error of exactly zero.
An error bar of 0.0 beside a score that is neither 0 nor 1 claims a rerun would return the identical number.
control 13,904 rows reproduce exactly, MATH does not
filed lm-evaluation-harness #3966
Swap one defensible aggregation choice, and fewer than half the #1s survive.
Same released data, a second reasonable way to combine it — recomputed and compared.
full top-5 holds 9 of 44
sharpest case changes no measurement at all
"We lead on subset X" was the luckiest of 23 noisy tests.
A "best-subset" win, corrected for how many subsets it could have been picked from.
Šidák-corrected (23 subsets) p = 0.19 (not significant)
a different, robust pair survives correction at p = 0.011
A "super-linear power law" that came free from a minus sign.
A reported scaling exponent for warning lead-time, tested against what the definition alone hands you.
constant-onset null (definition only) α = 1.47 / 1.07
Dyck, drop 1 high-leverage point α = 0.97, higher R²
The most-cited #1 in the field — I tried to break it and couldn't.
An independent Bradley-Terry recomputation of the public arena battles, bootstrapped for rank confidence.
P(truly #1) 1.00 (0.96 on coding votes)
one 19k shard P = 0.83, 3-way tie — volume is the story
The model shown second wins more — and it isn't the better model.
An order-bias check on the same public human votes, controlled within each matchup.
within-matchup, controlled for strength +1.25 pp, p ≈ 9×10⁻⁵
verdict real recency bias — randomization launders it, but only if you randomize
Pick the depth you need.
Every engagement ends in one plain-English readout: which claims hold, which don't, and the one-line statistical fix for each that doesn't. Fixed price, no retainer required, NDA on request.
- One benchmark / leaderboard / judge claim
- Reproduced from your data
- One-page verdict + the fix
- 72-hour turnaround
- Judge bias: position, verbosity, self-preference
- Multiple-comparisons & look-elsewhere
- Metric-artifact & averaged-away-bias checks
- Robustness: leave-one-out, mechanical nulls
- Written report + a call to walk it
- The 5-probe check as a CI gate
- Blocks a confounded metric on PR
- Wired to your eval pipeline
- Quarterly re-audit + tuning
Not ready to hand it over? Run the free browser checks first — eight of the checks I run in an audit, in your browser, on your own numbers. Nothing is uploaded.
Six questions I ask every number.
Not opinions — arithmetic. Each is a one-line check that a careful team could run itself; the value is that I run all of them, adversarially, on a claim you're too close to.
What would the definition give you for free?
Replace the interesting variable with a constant. If the "law" survives, it was in the algebra.
How many slices could you have picked?
Correct the "best subset" for the family it was chosen from. Most wins don't survive it.
Does one row flip the verdict?
Leave-one-out on every fit. A conclusion that hangs on a single high-leverage point isn't one.
Is the winner winning, or just longer?
Position swap, verbosity, self-preference. Certify the ranking, or expose the bias wearing its label.
Did the average fold the structure flat?
When mean-absolute equals mean-signed, "scatter" is a one-directional gap the summary hid.
Did the winner win, or get lucky?
Resample every entry from its own error bar and re-rank. If the crown moves, the ranking selected on noise — and the winning score is inflated by having been the maximum.
If it holds, I say so.
A tool that only ever finds fault is a cynic. Clean claims get a citable clean bill of health.
Measured, Not Believed
The book behind the practice — why AI benchmark scores and trading backtests overpromise, and how to catch them. Twenty-one chapters of real audits, from LLM judges to gravitational waves. $14.97 minimum, $29 suggested.
Send me the claim you're least sure about.
The one line in the paper, the slide, or the launch post you'd least want a reviewer to poke at. I'll tell you whether it holds — before anyone else looks.