← Research

91% accurate. Catches 46% of the failures.

36,063 fresh judgments from three 2026 open-weight models over JUDGE-BENCH's human-labelled datasets, measured as instruments rather than ranked as contestants. Accuracy hides the direction a judge is wrong in; the usual corrected-rate interval covers the truth 13–46% of the time; a calibration dies in transport.

•10 min read
91% accurate. Catches 46% of the failures.

36,063 fresh judgments from three 2026 open-weight models over JUDGE-BENCH's human-labelled datasets, re-measured as instruments rather than ranked as contestants. Every number in this post regenerates from the study's code. Companions in this series: The "95%" confidence interval that's right 13% of the time and A true 80/20 prints as 70/30.

An LLM judge marks your production traffic pass or fail, a pass rate goes on a dashboard, and somebody ships on it. The judge itself was validated — if at all — by an accuracy number on a benchmark.

This post asks what that number is worth, using the largest public collection of human-labelled judgments there is. We did not rank the judges. We measured them the way you would measure any diagnostic instrument: how often does it catch a real failure, how often does it flag a good answer, how well do you know each of those, and where does the answer stop being true.

Five findings. Each is a number with an interval.

Setup

JUDGE-BENCH (Bavaresco et al., ACL 2025): 20 datasets of human-labelled text under a shared schema — tens of thousands of instances spanning task types, quality dimensions, categorical and graded labels, and expert and non-expert annotators. Several datasets carry multiple annotations per item, which is what makes a human ceiling computable. We re-analysed the benchmark's published human annotations across its datasets, and generated the 36,063 fresh judgments on a battery of its binary-criteria sets — ToxicChat, CoLA, DICES-350 (expert and crowd labels), inferential-strategies, and ROSCOE/GSM8K.

Judges: three open-weight 2026 models, provider-direct, temperature 0, each dataset's own criterion and prompt, verbatim, per the benchmark's protocol.

The positive class is FAIL. A judge exists to catch bad outputs. Sensitivity is measured on the rows a human called bad.

No imputation. Unparseable output stays unparseable and is counted.

1. Accuracy hides which direction the judge is wrong in

ToxicChat: 2,853 real messages, 12.7% of them toxic.

judgeaccuracyshare of toxic messages caught
Llama-3.3-70B92.4%58.2%
Qwen3-30B-A3B92.0%80.6%
gpt-oss-120b91.2%46.4%
"not toxic" to everything87.3%0%

Three judges within 1.2 accuracy points of each other, and one catches nearly twice as many real failures as another. A constant that says nothing scores 87.3%.

This is not subtle statistics. It is what class imbalance does to a single-number metric, and it is why every number in this post is reported as a pair — TPR on the human-fail rows, TNR on the human-pass rows, each with a Wilson interval.

The operational consequence: these are different instruments for different jobs. A judge with high specificity and mediocre sensitivity is a precision instrument — when it flags something, believe it, but do not use it to tell you how much is broken. A judge with the opposite profile is a recall instrument — good for surfacing candidates for review, bad for gating. Accuracy cannot distinguish them, and the two failure modes have opposite costs.

2. A "95%" interval on a corrected rate covers the truth 13% of the time

Once you know a judge's sensitivity and specificity you can correct its observed pass rate for its own error — Rogan–Gladen, standard in diagnostic testing for decades. The correction is right. The interval usually is not.

The common implementation maps the observed rate's interval through the correction with TPR and TNR treated as exact constants. They are not exact. They are estimates from a finite calibration set, and their uncertainty never reaches the reported interval.

We tested this on real judge errors with real calibration splits: the naive corrected interval covers the truth 13–46% of the time while claiming 95%.

The failure has a nasty shape. Score more traffic and the interval converges — on a number whose real uncertainty is the several points the calibration set left in TPR/TNR. It looks tighter exactly as it becomes more wrong.

The fix is not new either. Lang and Reiczigel (2014) fold the Se/Sp estimation uncertainty into the interval and hold near-nominal coverage down to calibration sets of about 30. Measured here: 94–100% coverage, by being honestly wide.

And on a judge with no diagnostic power, the calibration-aware interval does something the naive one cannot: more calibration data teaches it to refuse rather than to emit confident garbage. Below a Youden's J floor, 1/J is an error amplifier — at J = 0.15, one point of judge error becomes nearly seven points of "corrected" rate. The honest output there is the observed rate, labelled uncorrected.

3. A κ target means nothing without the population's own ceiling

κ 0.31 is a failing grade on one dataset and human parity on another. The number is only interpretable against how much the humans agreed with each other.

DICES safety — 123 annotators per item. Human ceiling: κ ≈ 0.31. Every judge ever measured on it scores at or below zero, including all three 2026 models. There is a real signal in the consensus of 123 people that no single reader recovers.

Newsroom fluency — 3 annotators per item. Human ceiling: ρ ≈ 0.02. Eleven of twelve published judges score above it.

Both leaderboard readings are broken, in opposite directions. On DICES a judge is being asked to hit a target no single instrument reaches. On Newsroom, judges beating a near-zero ceiling are not good — they are agreeing with each other about a signal the humans do not share, which is shared bias, not accuracy.

The ceiling estimator, and why it needed one. Comparing a judge scored against a 3-rater mean to a leave-one-out human scored against a 2-rater mean is not a fair comparison: averaging three noisy ratings gives a less noisy target than averaging two, so the human's ceiling is artificially depressed. We derived a symmetric estimator,

c_k = r̂ / √(r̂ + (1 − r̂)/k)

and validated it against measured leave-one-out values before using it — WMT en-de: predicted 0.581 against 0.578 measured; the recipes dataset within ±0.02 on all six criteria.

The finding survives the correction. Newsroom fluency's fair ceiling is 0.00 — averaging three raters who share no signal does not create one — so eleven of twelve judges remain above it. And WMT en-de is the counter-case that shows the estimator is not simply producing zeros: fair ceiling 0.62, with only GPT-4o marginally above.

4. Well-posedness beats rubric specificity

The intuition is that vague criteria calibrate badly and specific ones calibrate well. That is not what the data says.

"Missing steps" — which reads as vague — reached Youden's J of 0.90. "Step grammar" — which reads as precise — broke completely, because unnatural phrasing has no stable meaning on a row like 3 + 4 = 7 duck eggs.

The variable is not how specific the wording is. It is whether the question is well-posed on the population being judged: whether a competent human, reading only that question and that row, would give the same answer twice.

That reframes what to do about a badly-agreeing criterion. Adding adjectives does not help. Splitting a criterion that is silently asking two questions does: stratify the disagreement on a traffic facet, and a criterion that is really asking about the agent's behaviour and about retrieval quality at the same time falls apart into two questions that each calibrate. That diagnosis — not more labels — is what a badly-agreeing criterion needs first.

5. A calibration is a (judge, criterion, population) triple, and it dies in transport

Reusing a judge's measured error rate on a different population is the failure mode with the worst consequences, because it is invisible and it produces a confident number.

Same source, different slice: 0.8 points of error. Tolerable.

Same criterion name, new population: a 36% reported failure rate where the truth was 0.6%. That is not a noisy estimate. It is worse than not correcting at all — the sophistication actively hurt, because it took a wrong observed rate and made it more wrong with an authoritative-looking interval attached.

And the ground truth moves too. On the same conversations, an expert annotator pool called 50% unsafe and a crowd pool called 77% unsafe, with κ = 0.31 between the two "truths". Which humans you ask is part of the population.

What we changed in errorbar

Everything above is enforced in code, not described in copy:

  • We never report accuracy or raw agreement. TPR and TNR with Wilson intervals, per direction, positive class = fail.
  • Every corrected rate ships with a Lang–Reiczigel interval carrying calibration uncertainty, with truncation counted and surfaced, and a calibrationVarianceShare that tells you when the interval is limited by how many rows you graded rather than how much traffic you scored.
  • Below the Youden floor we refuse to correct and say so.
  • A judge's trust verdict reads the interval against the bar, not the point estimate — and distinguishes "grade more, here is how many" from "grading will not settle this, narrow the question."
  • Calibrations are bound to a population at birth and voided when the scope, prompt, model, pre-stage or ground truth changes.
  • A judge's certificate prints the golden set's own human κ as the floor under every number on it. A judge cannot be claimed more reliable than the ground truth it was measured against.

We also shipped 33 starter judge cards built from this work. Seven of them are flagged "do not correct with this." That is the finding working, not the product failing.

Why we built this

Our own platform corrects production pass rates for judge error. We wanted to know what the correction was worth before we asked anyone to act on one.

The answer was that the method is sound and the standard implementation of the interval is not; that the metric everyone reports hides the direction that matters; and that a calibration is a perishable claim about a specific population rather than a property of a judge.

None of those are opinions about how evaluation should be done. They are measurements, on public data, with the code attached.

Limitations, stated

Three judges, one prompt per criterion, no prompt tuning — a differently-prompted judge would move these numbers, and we make no claim about the best achievable performance of any model here. JUDGE-BENCH's datasets vary widely in annotator count, expertise and label design; the ceiling analysis is only possible on the subset with multiple annotations per item, and for two-rater datasets those ceiling estimates carry wide intervals of their own. The interval and transport findings do not depend on judge quality — they would hold for a perfect judge measured on a finite sample — but the agreement figures are specific to these judges on these tasks. On DICES, all three judges score negative κ against both rater pools; we verified this is not a label-polarity error in our pipeline, and report it as measured.