← Research

The “95%” confidence interval that’s right 13% of the time

A re-analysis of JUDGE-BENCH plus 36,063 fresh judgments from three 2026 open-weight models. Five things a single agreement score hides — and what an LLM judge's error bars should actually look like.

•20 min read
The “95%” confidence interval that’s right 13% of the time

A re-analysis of JUDGE-BENCH, plus 36,063 fresh judgments from three 2026 open-weight models. Every number in this post regenerates from the study's code and raw judgments.

If you use an LLM judge to estimate a failure rate, correct for the judge's known error profile, and put error bars on the result the way most eval stacks do — binomial interval, calibration treated as known — your "95%" interval contains the truth as little as 13% of the time. That is measured, not simulated: real judges, real human labels, hundreds of real calibration splits. It is the sharpest of five findings in this study, and none of the five is visible in the number LLM-as-judge evals usually report — a single agreement score. This post is about what that one number hides. We took JUDGE-BENCH (Bavaresco et al., ACL 2025) — the best public collection of human-labeled judge data we know of: 20 datasets with genuine human annotations, from toxicity to translation quality — and asked five questions the leaderboard number can't answer:

  1. If you split "agreement" into catches real failures (TPR) and doesn't cry wolf (TNR), what do judges actually look like?
  2. If you use a judge to estimate a failure rate and correct for its known errors, how wrong is the usual confidence interval?
  3. How well do the humans behind the ground truth agree with each other — and what does that ceiling mean for judge scores?
  4. Do concrete rubrics make judges more measurable than vague ones?
  5. Does a judge's measured error profile survive a change of population?

The benchmark's aggregate results are published, but per-item judge outputs are not — so for questions 1, 2, 4, and 5 we generated fresh judgments with three current open-weight models (gpt-oss-120b, Llama-3.3-70B-Instruct, Qwen3-30B-A3B), using each dataset's own prompt, verbatim, per the benchmark's protocol (rendered template + "Answer with one of X, Y. Do not explain your answer.", temperature 0). Question 3 is pure re-analysis of the published human annotations. Total cost of all 36k judgments: single-digit dollars.

One methodological difference we should flag up front: the benchmark's harness assigns a random label when a judge's response doesn't parse (it tracks validity separately). We record unparseable as unparseable and report coverage instead. Genuine parse failures peaked at 0.6% of a cell (Llama-3.3-70B declining to label some of the nastiest ToxicChat messages — a refusal is itself a judge behavior worth tracking); the only larger gap is Qwen answering "Unsure" on 6.6% of DICES items, which is an answer the dataset's own prompt offers, not a parse failure. If you reuse the benchmark harness, know the random-imputation branch is there.

1. Accuracy is the wrong axis for a screening judge

ToxicChat is 2,853 real user messages, 12.7% of them toxic by human label. Here is what "how good is the judge?" looks like on it, one row per judge:

JudgeAccuracyκTPR (catches toxic)TNR (passes clean)
gpt-oss-120b91.2%0.5346.4% [41.3, 51.6]97.7% [97.1, 98.3]
Llama-3.3-70B92.4%0.6258.2% [53.1, 63.2]97.4% [96.7, 97.9]
Qwen3-30B-A3B92.0%0.6780.6% [76.2, 84.3]93.6% [92.6, 94.5]
Say "non-toxic" to everything87.3%0.00%100%

The 120B judge has the top-tier accuracy — and misses more than half the toxic content. The 30B judge — a quarter the size — catches 80% of it, at a slightly worse accuracy. And the do-nothing judge is four points behind the leader. When the failure you're screening for is rare, accuracy is mostly measuring the base rate; the two numbers that matter are TPR and TNR, and they need their own confidence intervals (we use Wilson throughout — the brackets above).

The general rule this table demonstrates: judges have asymmetric error profiles, and the asymmetry — not the average — determines what the judge is for. gpt-oss-120b is a precision instrument on this criterion (what it flags is almost certainly toxic); Qwen3-30B is a recall instrument (what it passes is probably clean). Same benchmark, same "accuracy", opposite deployment semantics.

2. The 95% interval that covers the truth 13% of the time

Suppose you've labeled a calibration sample, measured your judge's sensitivity and specificity on it, and now use the judge on live traffic with the standard Rogan–Gladen correction to estimate the true failure rate. What error bars do you put on that estimate?

The tempting method: take the binomial interval on the observed judge rate and push it through the correction formula, treating Se/Sp as known constants. The problem: they are not constants — they came from a finite sample, and the correction divides by (Se+Sp−1), so calibration noise is amplified, not averaged away.

We tested this on real data — no synthetic judges. For each judge × dataset, we repeatedly split the human-labeled items into a calibration sample (n = 25–200) and a "production stream" (the rest), computed both intervals, and checked whether each covered the stream's true rate. Coverage of a nominal 95% interval, 300 splits per cell:

Datasetn_calNaive coversLang–Reiczigel coversNaive widthLR width
ToxicChat (gpt-oss)2513%100%0.0370.277
ToxicChat (gpt-oss)10032%94%0.0530.221
ToxicChat (gpt-oss)20042%94%0.0490.162
CoLA (Qwen3-30B)2528%100%0.0940.558
CoLA (Qwen3-30B)20074%97%0.0990.191

(Full grid: 3 judges × 3 datasets × 4 calibration sizes in the study's results/interval_sim.json.)

The naive interval is not slightly optimistic. At realistic calibration sizes it covers the truth 13–46% of the time while claiming 95%. The Lang–Reiczigel interval (Lang & Reiczigel 2014), which folds the calibration sample's uncertainty into the variance, covers 94–100% — by being honestly wider. That width is not a defect; it is the actual amount you know. It also shrinks the way it should: quadruple your calibration labels and the LR width roughly halves, while the naive width barely moves because it was never measuring calibration uncertainty in the first place.

A width of 0.28 at n=25 should not read as "calibration is hopeless" — it reads as a price list. Here is what labels actually buy you on the ToxicChat judges (LR half-width, coverage ≥94% at every size):

Calibration labelsHalf-width of the corrected rateWhat it's good for
25±14–16 ptsknowing whether the judge works at all
50±12–14 ptsruling out catastrophe
100±9–11 ptsscreening-grade monitoring
200±6–8 ptscatching real regressions
400±4–6 ptstracking month-over-month drift
800±3–5 ptsabout as tight as this judge gets

This is the curve behind the alignment-set floors errorbar enforces (30 labels minimum, 100 to be comfortable): 30 is where the machinery starts telling you the truth about how little it knows; 100 is where the truth starts being useful; 200–400 is where it can hold a regression gate. The naive interval never tells you any of this, because its width barely responds to calibration size — it was never measuring calibration uncertainty in the first place.

Two edge behaviors worth knowing. On DICES safety — where every judge we ran is at or below chance (next section) — the correction machinery refuses to produce a number once the calibration sample is large enough to see that Youden's J ≈ 0 (the correction divides by it). With a tiny calibration sample the refusal can't trigger yet, and the LR interval instead returns width ≈ 1.0 — "the corrected rate is somewhere between 0 and 100%" — which is exactly the honest answer. More calibration data correctly taught the system to refuse a useless judge; the naive interval on those same cells covered the truth as little as 0% of the time (never above 67%) while remaining confidently narrow.

3. No judge can beat the humans' agreement with each other — and when one does, worry

Every judge score on a human-labeled benchmark has a ceiling: how well a human from the same pool would score if graded the same way. We computed that ceiling for every JUDGE-BENCH dataset with multiple annotators per item, two ways: the benchmark authors' method (sample a rater, compare to the consensus — which that rater voted in), and a leave-one-out version (compare to the consensus of the others), which is unbiased and harsher. Then we put every published judge score, plus our three fresh 2026 judges, against it.

The result is a two-sided indictment of naive leaderboard reading:

Side one: the unreachable ceiling. DICES-350 has 123 raters per item labeling conversation safety. The consensus is rock-stable; the raters genuinely disagree with each other (Krippendorff α = 0.16; a random rater vs the rest: κ ≈ 0.31). Against this target, the best of the paper's 11 judges scores κ = 0.013. Our three 2026 judges: −0.13, −0.21, −0.12. Not "worse than humans" — anti-correlated with the consensus. Two years of model progress moved safety judgment nowhere.

A negative κ should make you suspect the pipeline before the judges, so we verified the polarity three ways. First, there is no label-mapping layer to invert: gold and predictions are compared as the same 'Yes'/'No' strings, taken respectively from the dataset and from the model's own text (audited: in 90 of 90 sampled rows the parsed label appears verbatim in the raw response; zero inversions). Second, the inversion test: flipping our judges' predictions yields κ +0.13 to +0.32. One trap to note — on the exactly-balanced expert set (175/175), flip-κ equals −κ by algebraic identity, so the flipped values prove nothing there by themselves. What carries the weight is what the flipped reading would entail: our three judges beating every judge the benchmark authors measured with their own independent pipeline on the same data by an order of magnitude, with gpt-oss landing precisely at the human ceiling. The authors' results — 11 judges clustered at zero-or-below — corroborate the unflipped signs. Third, our judges agree with each other positively (κ ≈ 0.21), so a single-model parse inversion would show up as one judge anti-correlated with the other two; instead all three move together. The negative κ is a property of the judges: they carry a strong "this is safe" prior on exactly the conversations expert raters flag (gpt-oss answers "safe" on 97% of expert-unsafe items), which under a balanced gold lands below chance. But note what the ceiling does to interpretation: a judge scoring κ = 0.31 here would be at parity with a human rater. "κ = 0.31" is a failing grade on CoLA and a perfect score on DICES. A κ target that ignores the population's own agreement level is a category error.

Side two: the impossible overachievement. Newsroom's Fluency criterion has 3 raters per item who essentially do not agree with each other (α = 0.03). Yet 11 of 12 published judges score above the human-human ceiling — the best at ρ = 0.63.

Before interpreting that, an objection we should raise against our own comparison: it is asymmetric as first computed. The judge is scored against the mean of 3 raters; the leave-one-out human is scored against the mean of the other 2 — a noisier target — so the judge gets an easier target by construction, and "beating the ceiling" could in principle be that artifact. So we corrected for it. Estimate single-rater reliability r̂ directly (Spearman between two distinct raters per item), then project the ceiling an independent rater would achieve against the same k-rater mean the judge faces: c_k = r̂ / √(r̂ + (1−r̂)/k), the Spearman–Brown-style symmetric ceiling. The model checks itself: its predicted leave-one-out ceiling matches the measured one on every dataset (WMT en-de: predicted 0.581, measured 0.578; Newsroom Informativeness: 0.352 vs 0.378; recipes: within ±0.02 across all six criteria).

The correction moves nothing. Fluency's fair ceiling is 0.00 — measured single-rater reliability is ≈ 0, and averaging three raters who share no signal does not create one — so 11 of 12 judges remain above it. Newsroom Relevance: fair ceiling 0.18, still 10 of 12 above. wmt-human zh-en: fair ceiling 0.14, still 10 of 11 above. (Where humans do agree, the correction also behaves: WMT en-de's fair ceiling is 0.62 and only 1 of 11 judges — GPT-4o at 0.634 — sits marginally above it.) The over-ceiling scores are real, and what they mean is what they meant: agreeing with a noise-dominated aggregate better than its own raters do is evidence of fitting a shared bias component (verbosity, formality — things both LLMs and rating averages drift toward), not evidence of judge quality. The criterion, as labeled by this population, has no stable meaning to exceed.

Where the data allows a clean reading, modern judges do show real progress: on inferential-strategies (reasoning soundness), gpt-oss-120b scores κ = 0.75 where the paper's best 2024 judge (GPT-4o) scored 0.42 — though our other two judges land at 0.29 and 0.33, so the generation gap belongs to one model, not to 2026. (One honesty flag: that dataset's two raters agree perfectly on all 300 items — α = 1.0 exactly — which almost certainly means adjudicated consensus labels entered as duplicates, so its true ceiling is unmeasurable. The κ comparison across judge generations stands; "0.75 out of a possible 1.0" does not.)

The operational rule we take from this: an agreement target for a judge must be quoted relative to the measured human ceiling of its own labeled population, or it's not a target — it's a vibe.

4. A judge is calibratable when the question is well-posed on the population — not when the rubric merely sounds specific

Sort every (judge, criterion) pair we measured by Youden's J (TPR + TNR − 1 — the judge's usable signal above chance). J is what determines whether a judge can be calibrated at all: the rate correction divides by it.

CriterionBest JWorst JAll 3 judges usable (J > 0.05)?
Missing steps (ROSCOE)0.900.86yes
Toxicity (explicit guidelines)0.740.44yes
Final-answer detection0.740.67yes
Step repetition0.720.13yes
Reasoning soundness0.710.39yes
Contradiction0.660.59yes
Grammaticality (CoLA)0.650.40yes
Step hallucination0.620.44yes
Step grammar (ROSCOE)0.54−0.17no
Safety (DICES, expert labels)−0.15−0.26no
Safety (DICES, crowd labels)−0.20−0.26no

The floor of the table is the broad construct: "is this response safe" produces negative J for every judge against both rater pools — not weak, anti-correlated, and uncorrectable. No amount of calibration labeling fixes a judge whose question the raters themselves resolve by a different theory than the model does. The engineering consequence: decompose broad constructs into operational sub-criteria before calibrating (the safety-annotation literature has said this about DICES-style constructs all along — "safe" is a bundle of separately-ratable harms).

But the table's most instructive row is not the floor. It is "step grammar" — and it deserves more than a passing mention, because it breaks the heuristic everyone actually uses.

Wording specificity is not well-posedness

The advice "make your judge criteria specific" is standard, and by its lights "step grammar" is a model rubric: "Does this step contain any faulty, unconventional, or controversial grammar usage? In other words, does the language in this step sound unnatural?" Concrete words, single dimension, yes/no answer. It is also the worst-performing criterion outside safety — because the population it is asked about is math-reasoning steps like 3 + 4 = <<3+4=7>>7 duck eggs are used, and on that register "sounds unnatural" has no stable referent. Is calculator markup unnatural? Is telegraphic arithmetic prose "faulty grammar"? The rubric doesn't say, because it was written for language, and the population is math wearing language's clothes.

There is a clean way to see that this is instability of the question and not weakness of any model, and it needs no human labels at all: measure how well the three judges agree with each other, per criterion:

CriterionInter-judge κ (mean pairwise)
Missing steps0.92
Contradiction0.92
Final-answer detection0.76
Step hallucination0.71
Toxicity0.62–0.63
Grammaticality (CoLA)0.49
Reasoning soundness0.31
Step repetition0.27
Safety (DICES)0.21–0.22
Step grammar0.12

Three different model families, same rubric, same items — and on step grammar they agree with each other less than on anything else we measured, including safety. Each model resolved the ill-posed question with a different private theory of what it asks. Contrast "missing steps", which sounds holistic ("would adding steps make for a well-supported chain?") but is operationally checkable on GSM8K chains: judges agree with each other at κ = 0.92 and with humans at J = 0.90.

So the variable that predicts calibratability is not how specific the rubric sounds. It is whether the question has one answer-theory on the actual population — call it well-posedness. Specificity is a property of rubric text; well-posedness is a property of the rubric–population pair. You can estimate it before labeling anything: run two or three cheap judges over a sample and measure their agreement. Where they can't agree with each other, no calibration against humans will hold — and, as the next section shows, that is also exactly where a calibration measured elsewhere dies on arrival.

5. A calibration is a (judge, criterion, population) triple

The most seductive shortcut in judge deployment: measure Se/Sp once, keep using it. We tested transport in three regimes, same judge and criterion throughout:

RegimeSe cal → targetSp cal → targetTrue fail rateCorrected estimateError
Same source, held-out split (ToxicChat train→test, Qwen)0.79 → 0.810.94 → 0.9412.7%13.4%0.8 pts
Same construct, new population (CoLA → ROSCOE step-grammar, Qwen)0.79 → 0.200.86 → 0.630.6%36.4%35.9 pts
Same items, different rater pool (DICES expert → crowd)———refuses—

Act one: within a population, transport works — the correction cut the raw estimate's error from 3.1 points to 0.8. This is the case that makes people trust the shortcut.

Act two: same criterion name ("grammar"), different population (linguistics example sentences → math reasoning steps), and every judge's error profile moves. Qwen's sensitivity collapses from 0.79 to 0.20 and its transported correction turns a 0.6% true failure rate into a reported 36%; Llama lands at 35%. The corrected number is worse than not correcting at all. The third judge (gpt-oss) reports 0.0% — numerically fine, but only because its correction went negative and clamped: its calibrated specificity (0.52) was so far from the target's (0.94) that the formula overshot zero. Two catastrophes and one right-for-the-wrong-reason is not a regime you deploy into.

Act three: the same 350 conversations, expert rater pool vs crowd rater pool. The two "ground truths" call 50% vs 77% of the identical items unsafe, and agree with each other at κ = 0.31. There is no correction to transport because the target itself is population-relative — and the judges are too weak on either version for the correction to engage at all.

So: a Se/Sp pair is not a property of a judge. It's a property of a judge on a criterion on a population, and it expires when any leg moves. Measure on your traffic, re-measure when your traffic drifts.

What we changed in errorbar because of this

errorbar is a judge-calibration layer for LLM applications, and this study is the public version of arguments we had internally. Concretely, the product now: reports TPR/TNR with Wilson intervals instead of accuracy; uses the Lang–Reiczigel interval (the study's stats_lib.py is a line-for-line port of the production code) for every corrected rate; refuses to correct when Youden's J is below 0.15 rather than emitting confident garbage; binds every calibration to the population it was measured on and expires it after 30 days; and treats human-ceiling estimates as the reference point for κ targets. None of that is a sales pitch for anything in this post — every method here is standard statistics from the diagnostic-testing literature, and you can implement all of it in an afternoon.

Starter judges (clearly labeled: not your population)

The study's results/starter_judges.json ships 33 pre-measured judge cards — (model, criterion) pairs with TPR/TNR + Wilson intervals, κ, parse coverage, and a machine-readable usable_for_rate_correction flag — measured on the public populations above. 26 are usable for rate correction; 7 ship explicitly flagged "do not correct with this", because publishing a judge's uselessness on a construct is as much a deliverable as publishing its strength. They are starting points with honest error bars, not calibrations for your traffic; the transport section is the reason that caveat is printed on every card.

Method notes and honest caveats

  • Prompts: each dataset's own template, rendered per the benchmark's replace_instance, with the benchmark's "regular" answer-format suffix. One prompt per criterion; we did not tune prompts, sample multiple completions, or use CoT. Judges could do better with better prompting — this measures the benchmark's protocol, not the models' maximum.
  • Parsing: case-insensitive exact match on the first line, else first standalone label token; no random imputation. Coverage ≥ 99.4% per judge.
  • Binarization: graded 1–5 criteria binarize at fail ≤ 3 (both sides); categorical datasets used their native binary labels. DICES "Unsure" majorities excluded from binary analysis (counted, reported); "Unsure" judge answers treated as non-answers.
  • Ceiling estimators: sampled-rater κ/ρ, 300 simulations, both authors-style and leave-one-out. For 2-rater datasets the LOO ceiling is rater-vs-rater.
  • What we did not do: rank judges overall (three models, cherry-pickable), extend to other benchmarks, or claim the fresh judgments are comparable to the paper's leaderboard beyond the datasets we reran under the same protocol.
  • Population disclaimer, once more: every number here is a property of these datasets' populations. That is not a caveat about this study; it is the point of this study.