Measured, Not Believed — eval integrity live status → free checks →
Independent audit practice

Most benchmark wins don't survive a second look.

I find the ones that don't — before you publish them. Independent, reproducible audits of the numbers your model, your eval, or your leaderboard is about to stake its credibility on. If a claim is real, you get a clean bill of health you can cite. If it isn't, you find out from me, quietly, not from a reviewer or a competitor on launch day.


Case files · public data, reproduced

Eleven audits. Ten numbers that overstated — one that held.

Each was recomputed from public, released data. Ten target the one summary sentence that overpromised — a judge's bias, a lucky subset, a mechanical "law", an order effect, a base rate, a hidden second dimension, a crown won by luck. The eleventh is the control: a #1 I tried to break and couldn't, so I certify it. An auditor who only ever says no isn't checking anything. This is exactly the work I'd do on yours.

MT-Bench · GPT-4 judge

The judge preferred longer answers and its own family.

A widely-used LLM-as-judge, audited on its own released pairwise verdicts.

longer answer wins  68.0% of pairwise verdicts
judge's own model family wins  71.5%
both  p ≈ 0, fully reproducible from released data
bias, not skill Read the full write-up →
Base rates · clinical & security

0.978 AUC on the benchmark. About 12% precision in the clinic.

The same model, reprojected from a balanced test set to the prevalence an operator actually faces.

benchmark  AUC 0.978 on a balanced set
at 1% prevalence  ≈12% precision — 7 of 8 alarms false
at 0.1%  ≈1%, deterministic from Bayes
different world, same model Read the full write-up →
MTEB · ProteinGym · SWE-bench

The ranking is one column. The data has up to six dimensions.

Parallel analysis on the model × item matrix, testing whether "better" is really a direction.

MTEB  k ≈ 6 independent axes
ProteinGym  k ≈ 3 — second axis has a nameable mechanism
SWE-bench  k ≈ 2, each gated by a permutation null
averaged away Read the full write-up →
7 boards · SWE-bench ×5, AlpacaEval 2, Aider

On five of seven boards, the announced #1 is more likely wrong than right.

Not a defect anyone can fix — sorting on score picks the largest error along with the largest ability.

SWE-bench Verified  72.1% chance the #1 is not the best
winning score  inflated +0.97 to +2.66 points
control  0.0% on the one board separated by 8 SE
the luckiest, not the best Read the full write-up →
HELM classic · MMLU college chemistry

Answering "D" to everything outranks 22 of the 29 published models.

The answer key is not uniform. A constant answer reads nothing and still lands above the board median.

always D  0.4100  ·  board median 0.2633
below chance  11 of 29 published runs
bound  same key on a later board beats only 13%
the roster, not the benchmark Read the full write-up →
Open LLM Leaderboard v2 · reported upstream

423 published records carry a standard error of exactly zero.

An error bar of 0.0 beside a score that is neither 0 nor 1 claims a rerun would return the identical number.

affected  107 of 420 models sampled
control  13,904 rows reproduce exactly, MATH does not
filed  lm-evaluation-harness #3966
false precision, still published Read the full write-up →
5 boards · 44 board/choice pairs

Swap one defensible aggregation choice, and fewer than half the #1s survive.

Same released data, a second reasonable way to combine it — recomputed and compared.

#1 holds  19 of 44 pairs
full top-5 holds  9 of 44
sharpest case  changes no measurement at all
the aggregation, not the models Read the full write-up →
RewardBench · 23 subsets

"We lead on subset X" was the luckiest of 23 noisy tests.

A "best-subset" win, corrected for how many subsets it could have been picked from.

raw  p = 0.009  (looks significant)
Šidák-corrected (23 subsets)  p = 0.19  (not significant)
a different, robust pair  survives correction at p = 0.011
max of many · not a finding Read the full write-up →
Grokking · loss-landscape geometry

A "super-linear power law" that came free from a minus sign.

A reported scaling exponent for warning lead-time, tested against what the definition alone hands you.

reported  α = 1.18 (SCAN), 1.13 (Dyck) — "super-linear"
constant-onset null (definition only)  α = 1.47 / 1.07
Dyck, drop 1 high-leverage point  α = 0.97, higher R²
mechanical · fragile Read the full write-up →
LMArena · the control · certified

The most-cited #1 in the field — I tried to break it and couldn't.

An independent Bradley-Terry recomputation of the public arena battles, bootstrapped for rank confidence.

gemini-2.5-pro over 98,088 battles  rank CI [1, 1]
P(truly #1)  1.00  (0.96 on coding votes)
one 19k shard  P = 0.83, 3-way tie — volume is the story
resolved · certified Read the full write-up →
LMArena · human-vote order bias

The model shown second wins more — and it isn't the better model.

An order-bias check on the same public human votes, controlled within each matchup.

second-shown model wins  50.62%  vs 49.38% (98k votes)
within-matchup, controlled for strength  +1.25 pp, p ≈ 9×10⁻⁵
verdict  real recency bias — randomization launders it, but only if you randomize
order effect · real Read the full write-up →

The engagement

Pick the depth you need.

Every engagement ends in one plain-English readout: which claims hold, which don't, and the one-line statistical fix for each that doesn't. Fixed price, no retainer required, NDA on request.

Spot Check
$600
One claim you're about to publish — a subset win, an exponent, a "remarkable agreement".
  • One benchmark / leaderboard / judge claim
  • Reproduced from your data
  • One-page verdict + the fix
  • 72-hour turnaround
Book Spot Check →
most booked
Full Eval Audit
$3,500
Your eval or leaderboard, end to end, before a launch or a paper.
  • Judge bias: position, verbosity, self-preference
  • Multiple-comparisons & look-elsewhere
  • Metric-artifact & averaged-away-bias checks
  • Robustness: leave-one-out, mechanical nulls
  • Written report + a call to walk it
Book Full Audit →
CI Gate
from $1,000 /mo
Stop a confounded number before it ships — every release, automatically.
  • The 5-probe check as a CI gate
  • Blocks a confounded metric on PR
  • Wired to your eval pipeline
  • Quarterly re-audit + tuning
Start CI Gate →

Not ready to hand it over? Run the free browser checks first — eight of the checks I run in an audit, in your browser, on your own numbers. Nothing is uploaded.


Method

Six questions I ask every number.

Not opinions — arithmetic. Each is a one-line check that a careful team could run itself; the value is that I run all of them, adversarially, on a claim you're too close to.

01 · before

What would the definition give you for free?

Replace the interesting variable with a constant. If the "law" survives, it was in the algebra.

02 · look-elsewhere

How many slices could you have picked?

Correct the "best subset" for the family it was chosen from. Most wins don't survive it.

03 · fragility

Does one row flip the verdict?

Leave-one-out on every fit. A conclusion that hangs on a single high-leverage point isn't one.

04 · the judge

Is the winner winning, or just longer?

Position swap, verbosity, self-preference. Certify the ranking, or expose the bias wearing its label.

05 · the tell

Did the average fold the structure flat?

When mean-absolute equals mean-signed, "scatter" is a one-directional gap the summary hid.

06 · selection

Did the winner win, or get lucky?

Resample every entry from its own error bar and re-rank. If the crown moves, the ranking selected on noise — and the winning score is inflated by having been the maximum.

→ and

If it holds, I say so.

A tool that only ever finds fault is a cynic. Clean claims get a citable clean bill of health.

Run all eight checks free in your browser →


The credential

Measured, Not Believed

The book behind the practice — why AI benchmark scores and trading backtests overpromise, and how to catch them. Twenty-one chapters of real audits, from LLM judges to gravitational waves. $14.97 minimum, $29 suggested.

About the book →
Book an audit

Send me the claim you're least sure about.

The one line in the paper, the slide, or the launch post you'd least want a reviewer to poke at. I'll tell you whether it holds — before anyone else looks.