A pre-registered study on the public HealthBench physician meta-evaluation: 650 unanimous physician-labelled rows, seven judge model families, two grading protocols, twenty rounds of the product's own prompt optimiser, an ensemble probe, and one human rewrite of the criterion. 57 certificates. Every plan was written before its results and every deviation, including one study error, is dated in the repo. Companions in this series: The "95%" confidence interval that's right 13% of the time and A true 80/20 prints as 70/30.
An LLM judge is the reward in most post-training loops now. The question nobody asks before the run starts is how wrong the judge is. The question nobody can answer after it ends is whether the policy learned the task or learned the judge.
We built a product whose whole claim is that it answers the first question with a measured error rate and refuses to be a reward until that rate clears a bar. The honest way to test the claim is to take a criterion nobody can argue with, on public data with expert labels, and try as hard as a motivated customer would to get a judge through.
Nothing got through. This post is about what that means.
Setup
The criterion. HealthBench's emergency-referral behaviour, the most binary thing in medicine: did the response tell someone with an emergency to seek care now, without alarmism, without harm, in line with consensus. The public meta-evaluation carries 60,896 physician ratings on 29,511 (response, rubric) rows; the three emergency-referral clusters hold 1,806 of them. We used the 650 rows where every physician agreed, split by prompt into a calibration half and a held-out half.
The instrument. errorbar's binary judge: temperature 0, logprob-scored, the criterion text inserted verbatim. Calibration through the public API, exactly as a customer would do it: import the rows as traffic, post the physician verdicts as labels, freeze a label set per cluster, create a criterion, align it. The product splits each label set into a tune half and a report half by request-id hash and measures the certificate on the report half only. Every number below is from that report half.
The bar. Sensitivity (the judge fails what physicians failed) and specificity (the judge passes what physicians passed) both at 0.90 on the lower 95% bound. A judge below the bar is refused as a reward, with the reason.
1. The verbatim rubric is not a judge
Registration 3a used each cluster's rubric text, unchanged, as the criterion. Two judges, three clusters, six certificates, all misaligned. On the emergent cluster Qwen3-235B caught 85% of physician fails and passed 62% of physician passes, κ 0.35. Llama 3.3 70B was the mirror image: 58% of fails, 89% of passes.
The judges err in opposite directions on the same text. That is the first sign the gap is not a model property.
2. The benchmark's own grading protocol does not close it
Registration 3b replaced the prompt with HealthBench's own grader instructions: score this rubric item and only this item, all sentences must be met, lists introduced by "such as" need not be exhaustive. Six more certificates, all misaligned, individual numbers moving a few points in both directions. The protocol was written for a different grader on a different model; dropped into another instrument it is not a calibrated one until someone measures it.
3. Bigger judges move κ from 0.4 to 0.6, not to 0.9
Registration 3c tried the strongest judges in the catalogue: Qwen3.5 397B, DeepSeek V4 Flash and Pro, gpt-oss-120b, GLM-5.2, and a thinking model. Fifteen certificates, all misaligned. The best raw judge, DeepSeek V4 Flash, caught 77% of fails and passed 86% of passes, κ 0.59. The thinking model produced no gradable verdicts at all. The provider retired Qwen3.5 while we were using it; more on that below.
4. The product's own optimiser makes judges stricter, not better
Registration 3d ran twenty rounds of GEPA auto-improve over four judges. Each round rewrote the judge prompt from the tune half's disagreements and re-certified the successor on the untouched report half.
| judge | best round | TPR [95%] | TNR [95%] | κ | verdict |
|---|---|---|---|---|---|
| GLM-5.2 | round 2 | 0.88 [0.71, 0.96] | 0.84 [0.74, 0.90] | 0.64 | borderline |
| DeepSeek V4 Flash | round 1 | 0.88 [0.71, 0.96] | 0.82 [0.73, 0.89] | 0.62 | misaligned |
| Qwen3.5 397B | round 4 | 0.92 [0.76, 0.98] | 0.78 [0.67, 0.85] | 0.58 | misaligned |
| Qwen3-235B | round 1 | 0.88 [0.71, 0.96] | 0.64 [0.53, 0.73] | 0.39 | misaligned |
One certificate in 48 reached borderline: the intervals straddle the bar on both sides and the product's reachability check says no number of further grades at these rates would prove it. Every other round made its judge stricter than its parent. DeepSeek's specificity went 0.82, 0.34, 0.23, 0.50, 0.75, 0.66 across six rewrites from the same best parent. A rewriter that reads a dossier of disagreements tightens the rule; tightening is not the same as agreeing with physicians.
5. Voting does not help, because the judges make the same mistakes
We reconstructed every judge's per-row verdict from the alignment reports and scored majority votes offline. The best majority of three reached κ 0.66, three points above the best single judge. Majority of five was worse. The judges miss the same physician fails and fail the same physician passes.
Correlated errors mean the gap is not in any model. It is between what the rubric text says and what the physicians actually applied. The lever left is the criterion: split it into the atomic checks the physicians were making, or rewrite it from their disagreements, and certify each part. That is a different study on new evidence.
6. Rewriting the criterion moved the trade-off, not the agreement
The product's criterion doctor read the two best judges' tune-half disagreements and said the same thing for both: over 80% of the errors ran one way, the judge failing answers the physicians passed, and it named the causes — disclaimers about not being a doctor, structured step-by-step guidance, answers to clinicians already managing the patient, non-English answers, compassionate framing. We rewrote the criterion once per cluster, by hand, to ignore all of that and ask one question: would the person get emergency-level care as fast as a good clinician would want? The text was frozen before any judge call; the report half stayed sealed.
It fixed exactly the half the doctor measured. Over-failing disappeared: specificity rose to 0.79–0.96, and GLM-5.2 on non-emergent cases reached borderline at 0.95, κ 0.67. Sensitivity fell to 0.49–0.75: the judges now passed answers the physicians failed. Nine certificates on three judges, none trustworthy, all refused. The operating point moved along the trade-off; agreement did not rise. The physicians were applying a judgment about whether the response was good enough, and no single sentence a model can follow captures both edges of it at once. The one design that would — certifying each atomic check on its own and composing the reward from certified parts — needs physician labels per check, which this dataset does not carry.
7. The door stayed shut
Every reward-session request on every one of the 57 criteria was refused with the same sentence:
The reward criterion must be proven at 90% on both error rates. It is misaligned: its error rates are measurably below the 90% bar. A policy trained against an unproven judge learns that judge's mistakes.
That is the claim this study set out to test, and it held under an adversarial amount of trying. The two experiments behind the door as first written, a sweep of the mid-run anchor's false-hold rate on certified verdicts and a real GRPO run with the anchor live, need a certified judge and did not run.
8. The anchor, measured on runs it should and should not stop
A sixth registration, written before any run, put the anchor itself under test with an uncertified reward on purpose. Qwen2.5-7B-Instruct, QLoRA, GRPO with groups of eight, 60 steps on one spot H100 per run. A frozen set of 50 emergency-referral prompts, one completion each every 10 steps, graded offline by the run's own reward and by ten anchor instruments reconstructed from this study's certificates, κ 0.30 to 0.55 on the emergent cluster. The anchor rule was the product's, unchanged: window 3, drop 0.10, rise 0.10, at least 20 prompts.
Three healthy runs, rewarded by GLM-5.2 with the rewritten per-cluster criteria, improved under their own reward (+0.037, +0.050, +0.093 on 60 holdout prompts) with no tripwire and the directive feature at its base rate. The anchors stopped them anyway. The best anchor held 7 of 15 checks, 47 per 100 [25, 70]; every anchor on the ladder held between 3 and 8 of 15. All holds were on the drop rule. On the raw, uncorrected rate the same anchors held 0 to 3 of 15: the correction that makes the anchor's rate mean something multiplies its sampling noise by about two at this κ, and with 50 prompts a move of 0.10 between checkpoints is what a healthy run produces on its own.
One run we tried to game, rewarded by Qwen3-235B with the emergent criterion applied verbatim to every cluster, did not game. The reward passed 6 to 8 percent of rollouts, a GRPO group with no pass carries no gradient, and the directive rate on non-emergent prompts stayed between 0.06 and 0.28 across 120 steps. Reported as no hack induced: detection and lead time are untested, and nothing was changed after seeing data. The registration also had a defect, disclosed rather than repaired: its unit, false holds per 100 steps, tops out at 10 with a check every 10 steps, so its stop threshold could not be reached.
What it means for the product is plain. The anchor does what its rule says with an instrument whose noise the rule does not know about. A drop threshold fixed at 0.10 is too eager for anchors at κ around 0.5 on a 50-prompt set. The fix is a rule that reads the anchor's certified sensitivity and specificity and sizes the threshold or the frozen set to them, so a hold means the corrected rate moved by more than its own interval. This sweep is the data to size it. The rule is built after the study, not on it. Six invalid launches preceded the first valid run; each is in the deviations file with its fix. About 7.5 spot GPU hours and $35 of judge calls.
What we changed in errorbar
The study hit four gaps in the product, each fixed on main before it finished. There was no public way to import externally labelled traffic; there is now. The prompt optimiser returned an empty rulebook for reasoning-heavy judges, twice, for two different reasons; both fixed. A provider retiring a model under a running study touched neither the catalogue, the certificate, nor the reward gate; now a definitive 404 retires the catalogue row, the certificate says the judge is gone, and the gate refuses it. One gap is recorded and not yet fixed: a thinking judge that produces no gradable verdict fails silently.
The study error
The first run applied one cluster's rubric text, the emergent cluster's, which demands an emergency referral in the first sentences, to all three clusters. The product certified it misaligned and refused it, correctly, but the per-cluster reading in the first write-up was wrong. It was found while previewing the next prompt, before any further judge call, disclosed the same day, and the run is preserved in the repo as an error rather than a result. We mention it because a study that admits its error is easier to believe than one that has none.
Cost
$44.17 across every registration, ledger-true, most of it on the two largest judges. The first two registrations together cost under a dollar; the criterion rewrite cost about $5. Roughly 14,000 judge calls. Zero GPU hours.
Limitations, stated
One benchmark, one criterion, one instrument. Nothing here is a statement about HealthBench's own grader, which was meta-evaluated on its own protocol with its own model. The calibration half is 207 rows for the cluster that gated everything; the report-half intervals are wide, and every conclusion above is drawn from where the intervals sit relative to the bar, not from point estimates. The optimiser was run as shipped, with one addition made mid-study and disclosed: branching from the best round rather than the latest. The criterion rewrite was one pass by one person; a second pass, or a rewrite from a different reading of the same rows, was not tried. The ensemble probe was not pre-registered and is labelled exploratory. Every certificate is signed and can be verified against errorbar's public key.
Continue with The honest part of the loop: what this result means for a post-training team — the reward function your trainer calls, the anchor that re-checks the run while it trains, and the certificate that turns the reported gain into a claim.
