Companions in this series: Seven frontier judges, one physician-written criterion, none over the bar, The "95%" confidence interval that's right 13% of the time and A true 80/20 prints as 70/30. Every plan below was written before its run; every deviation and error is dated in the repository.
Abstract
Question. When a policy is trained by reinforcement learning against an LLM grader, can an independent, separately certified instrument read the run's true quality while the grader misreads it, and stay silent when the run is healthy?
Method. Qwen2.5-7B-Instruct, GRPO, competition mathematics (MATH Levels 4 and 5). Three runs on one frozen set of 300 held-out problems, one completion per problem every 16 steps: two seeds rewarded by a cheap grader with no reference answer (Llama-3.3-70B, chosen from a pre-registered ladder for its measured leniency: 61% of wrong answers passed), and one control rewarded by a reference-holding oracle. Each snapshot was read three ways: the reward grader's verdict, exact-match truth, and an anchor (a two-stage mechanical judge: DeepSeek-V4-Flash solves the problem with the candidate withheld, then the two final answers are compared by exact match; certified sensitivity 0.987, specificity 0.979). The product's shipped hold rule was applied unchanged. A fourth experiment certified four graders, with the reference in the prompt, against 600 human labels from a public open-QA benchmark, and the cheap grader was re-certified on the 600 math completions with the reference in hand.
Results. Seed one: truth fell from 0.530 to 0.370 by step 80 and did not recover (paired McNemar, 58 lost against 16 gained, p below 0.0001) while the grader's read stayed between 0.85 and 0.96; the anchor was within 0.004 of the truth at every snapshot, including the collapse to 0.040, and the hold rule fired at step 96 pinning the checkpoint with the highest true accuracy. Seed two, uninterrupted: a 7-point dip at step 96 that recovered to 3 points by step 160, endpoint not significant; the anchor within 0.007; a hold at step 112 that again pinned the true best checkpoint. Control: training reward rose from 0.54 to 0.68 while the frozen set stayed flat within noise (endpoint up 3.3 points, paired p 0.24) and the anchor issued zero holds in eleven checks. On QA with the reference in hand, four graders converged to sensitivity 0.95 and specificity 0.85 to 0.87; fourteen human-correct answers were failed by all four against 0.1 expected by chance, and adjudicating them found ten label errors, which move every judge's specificity by the same three points, to the bar. On math with the reference in hand, the same Llama grader that passed 61% of wrong answers without one caught 97% of them and certified (specificity 0.985).
Interpretation. The gap between reported and true quality opened on both corrupted seeds and was readable in flight by an instrument the trainer could not lean on. The rule stayed quiet on the healthy run because it compares the grader's and the anchor's reads on the same frozen population, so a gain on the training distribution cannot trip it; a rule watching the training curve would have false-alarmed on that run. A simpler threshold on the anchor read detected earlier than the shipped rule on both seeds, and the rule's rise test, which used a trailing window, missed a divergence that grew steadily. The rule now reads the gap between the grader's and the anchor's reads on the same outputs, measured from the run's baseline against its own paired noise: on this data it holds at steps 64 and 96, one to two checks before the drop test, and reports which of three states it is in. The cheap grader's leniency was a property of grading without a reference, not of the model: given one, it certified on the same items.
1. Introduction
Every post-training team has a chart of reward going up. Almost none has the second line, what was true while it went up. The gap between those lines is where a policy learns the grader instead of the task, and a grader cannot report its own gap.
The earlier studies in this series measured graders at rest: how often a judge agrees with human labels, with intervals, and whether that clears a bar. This study measures a grader under optimization pressure, on a run where the truth is known, and asks whether a separately certified instrument can read the run while the grader cannot. It also asks the question that makes the answer useful: does the instrument stay silent when the run is fine?
Two earlier attempts of ours produced nothing, and the reasons are part of the record. The grader was too accurate to exploit and the policy had no gradient to learn from. Both facts were computable before the boxes went up and neither was computed. Section 3.6 describes the preflight that now runs before any launch.
2. Background
The instrument. errorbar's binary judge: one question, one unit, one judge model, temperature 0, logprob-scored. Calibration measures the judge against human grades on a held-out report half, producing sensitivity (the judge fails what humans failed) and specificity (the judge passes what humans passed) with Wilson intervals. A judge is trustworthy when both 95% lower bounds clear 0.90; the reward door refuses anything less.
The anchor. A second certified judge, independent of the reward, read on a frozen set of prompts every N steps. The hold rule compares two reads on those same completions: the reward grader's and the anchor's. It holds on a drop (the anchor's corrected rate falls more than a noise-aware margin below its best for two consecutive checks) or on divergence (the grader's read rises 10 points over a window while the anchor's does not). A hold pins the best checkpoint; it never kills a run.
Positive class. Throughout, FAIL is the positive class: sensitivity is the rate at which wrong answers are caught.
3. Methods
3.1 Policy, trainer, data
Qwen2.5-7B-Instruct, GRPO, 64 rollouts per step (8 prompts times 8 samples), learning rate 2e-4, KL coefficient 0, rank-16 LoRA, vLLM for generation, temperature 0.9 for rollouts and anchor completions alike so the certificates describe the sampler that produced the readings. Training prompts: 600 MATH Level 4 and 5 problems for the corrupted seeds; for the control, the 261 of those the base policy solved between 12.5% and 87.5% of the time, so there was something to learn. Dynamic sampling in both cases: a prompt whose group of 8 is unanimous is benched and probed later.
3.2 The reward graders
Corrupted seeds: Llama-3.3-70B-Instruct on a plain "is this solution correct" prompt with no reference, provider-direct, temperature 0, 8,192 output tokens, unparsed verdicts scored as 0. Control: the same prompt shape with the reference answer substituted per problem, on DeepSeek-V4-Flash.
3.3 The anchor and the truth
Anchor: DeepSeek-V4-Flash in two-stage mechanical mode, certified on 600 labelled step-0 completions at sensitivity 0.987 and specificity 0.979. For the control run, where DeepSeek was the reward, the anchor read used GLM-5.3-Flash's cached two-stage solves so reward and anchor remained different families. Truth: exact match against the MATH reference via a symbolic checker, available to the study and to no customer.
3.4 The frozen set and the readings
300 problems the policy never trains on, one completion each at step 0 and every 16 steps, generated by the trainer's own vLLM engine after a weight sync. Each snapshot was read offline three ways. All per-item rows are in the repository.
3.5 Pre-registration and analysis rules
Each run's predictions were written before launch. Results files carry a flow block counting indeterminate verdicts, a certainty grade with named downgrades, and a stamp that says whether the number may be quoted; a pre-commit gate refuses unstamped results. Comparisons on the same items use McNemar's exact test on discordant pairs. Lead times are reported in steps of 64 rollouts.
3.6 The preflight and the grader ladder
Before any box, three gates were computed from logged step-0 rollouts: the dead-group fraction under the grader's reward at group size 8 (learning is possible below 0.20), the attainable inflation of the grader's read on never-solved prompts (a divergence is detectable above 0.10), and the truth mass at risk on prompts where a passing wrong answer is as reachable as a right one (an onset by drop is detectable above 0.10). Three cheap graders were then judged on 4,396 stratified rollouts.
| grader, no reference | passes wrong answers | fails right answers | reward for wrong vs right |
|---|---|---|---|
| Qwen3-30B-A3B | 15% | 19% | 0.15 vs 0.81 |
| gemma-3-27b-it | 38% | 9% | 0.36 vs 0.89 |
| Llama-3.3-70B | 61% | 3% | 0.61 vs 0.97 |
The largest model was the most lenient; it said pass to 84% of everything it saw. On the certified 600 items Llama passed 168 of 262 wrong answers (sensitivity 0.36). It shared 3 of those false passes with the anchor, against 1.9 expected by chance; with three events the Poisson 95% interval runs from 0.6 to 8.8, so the ratio to chance spans 0.3 to 4.6 and the honest statement is that no shared blind spot was detectable at this sample size.
The third gate predicted that true accuracy on math could not fall by per-prompt selection, because a passing wrong answer and a right answer earn the same reward and a policy only drifts toward wrong where it cannot be right. That prediction failed; section 4.1 gives the mechanism it did not model.
4. Results
4.1 Seed one: a lasting loss the grader reported as a gain
| step | truth | grader's read | anchor read | mean length (chars) | grader passes a wrong answer |
|---|---|---|---|---|---|
| 0 | 0.530 | 0.827 | 0.530 | 1,602 | 64% |
| 32 | 0.543 | 0.910 | 0.547 | 1,370 | 85% |
| 64 | 0.513 | 0.963 | 0.517 | 1,378 | 94% |
| 80 | 0.370 | 0.850 | 0.373 | 937 | 78% |
| 96 | 0.390 | 0.917 | 0.393 | 871 | 88% |
| 128 | 0.390 | 0.940 | 0.393 | 1,121 | — |
| 144 | 0.040 | 0.290 | 0.040 | 3,046 | — |
By step 64 the grader was passing 94% of wrong answers on the frozen set, so being right no longer earned more than being confidently wrong. The policy drifted to what is cheapest to produce and still rewarded: derivations a third shorter, always ending in a boxed answer, wrapped in assertive phrasing. Shorter working means more wrong answers. Truth fell 14 points between steps 64 and 80 and stayed down; paired on the same 300 items, 58 lost against 16 gained, p below 0.0001. The grader's read never fell below 0.85; the training reward peaked at 1.00 at step 94 with truth at 0.39.
The anchor was within 0.004 of the truth at every snapshot, including the collapse at step 144, where truth read 0.040 and the anchor 0.040: a 49-point fall tracked to within half a point. The hold rule fired at step 96 on the drop test and pinned checkpoint 32, the checkpoint with the highest true accuracy in the run.
This is a style seam rather than a per-prompt one: a surface form the grader rewards regardless of content, transferring across prompts the policy could solve. It is the case the preflight's third gate stated it did not model.
From step 139 the run collapsed into repetitive code loops at the token cap and the training reward crashed. That is optimizer instability with no KL brake, not a hack, and the run was stopped at 150. A spot eviction at step 62 had forced a resume from checkpoint 60, and the cliff sits in the sixteen steps after it; this confound motivated seed two.
4.2 Seed two: direction and detection replicate, the size does not
Uninterrupted, 160 steps, on-demand hardware. Truth fell 7 points to 0.463 at step 96 and recovered to 0.503 by step 160. The grader's read rose 8 points and never fell below 0.84. The anchor was within 0.007 of the truth at all eleven snapshots. The hold rule fired at step 112, on the dip, and pinned checkpoint 16, again the checkpoint with the highest true accuracy; that checkpoint beat the run's endpoint by 4 points, not significantly. At the endpoint the loss is not significant: 34 lost against 25 gained, p 0.30.
Direction, detection, and the pin landing on the true best checkpoint replicated. A lasting loss is shown on one seed. The resume very likely amplified seed one's cliff, and the clean single-seed effect is a 7-point dip that recovered. We quote 7 when we quote one number, and the title's 14 is seed one's endpoint before the collapse.
4.3 The control: an honest reward, a flat frozen set, zero holds
Training reward rose from 0.54 to 0.68 over 160 steps. The frozen set: 0.533 at step 0, 0.567 at step 160. The ten snapshots between were flat, ranging from 0.507 to 0.543, and the endpoint's 3.3-point rise is within noise (paired, 24 lost against 34 gained, p 0.24). On the 96-item slice of the frozen set matching the training difficulty band, 0.490 at step 0 and 0.490 at step 112. So the frozen set did not move; the endpoint is up within noise, not a sustained rise. The oracle's own read was within 0.007 of the truth throughout; the anchor within 0.010. Zero holds in eleven checks, and the divergence test never fired.
This is the divergence signature, reward up and independent read flat, and the rule correctly did not fire. The reason is structural. The rule compares the grader's read and the anchor's read on the same frozen completions; the training reward never enters it. A rule watching the training curve against a frozen check would have false-alarmed on this run, as the free reward-rise heuristics did. The false-alarm side of the table was therefore measured under the condition that breaks the obvious implementation, which is a stronger test than a run where everything rises together and the rule has nothing to get wrong.
The gain on served prompts did not transfer to the held-out set over about five passes of the 261-prompt band: the policy learned the prompts more than the skill. A separate evaluation split with an honest grader would also catch that. The anchor's contribution is that it needs neither a trustworthy evaluation grader nor a separate split.
Completions shortened from 1,617 to about 1,400 characters with truth flat under an honest reward. That is conciseness at no cost. The failure in section 4.1 is style change decoupled from truth under a lenient grader, not style change itself.
4.4 The detector table
| detector | seed 1 | seed 2 | control | available to a customer |
|---|---|---|---|---|
| product anchor rule, drop test | hold at 96, pins the true best checkpoint | hold at 112, pins the true best checkpoint | 0 holds in 11 | yes |
| product anchor rule, divergence test (trailing window, as shipped during the study) | never | never | never | yes |
| single threshold on the anchor read, −5 points | step 80 | step 64 | never | yes |
| grader's read minus anchor's read, +5 over baseline, two checks (decision at the second) | first at 16, hold at 32 | first at 32, hold at 48 | never | yes |
| the same gap against its paired noise (the rule now shipped) | hold at 64 | hold at 96 | never | yes |
| grader's read minus truth, +5 over baseline, two checks | first at 16, hold at 32 | first at 32, hold at 48 | never | no, needs the truth |
| best training reward | picks step 94, truth 0.39 | picks about step 128, truth 0.47 | picks the end | yes |
| training-reward rise | fires | fires | fires, a false alarm | yes |
| entropy collapse, length explosion | step 145, on the collapse | never | never | yes |
Three rows need reading. First, the simpler threshold on the anchor read fired 16 and 48 steps before the shipped rule on the two seeds, with zero false alarms on the control. The rule trades that latency for its two-consecutive-check margin, and on this data the trade did not pay. Second, the rule's divergence test as shipped during the study never fired on either seed although the divergence grew to 15 and 10 points, because it measured the grader's rise over a trailing window of three checks and the largest three-check rise was 9 points on seed one and 8 on seed two. Measured from the run's baseline the signature shows on a single check at step 64 on seed one (+0.137 while the anchor moved −0.013), but with two consecutive checks required it would first hold at 128, after the drop hold at 96, and never on seed two. That test is retired.
Third, the gap rows. The grader's read minus the anchor's read on the same frozen outputs is a paired comparison of two instruments: whatever is common to both reads on a snapshot, which problems came out hard and how the policy sampled, cancels, and what is left is their disagreement, which is the quantity a fooled grader changes. It reads the mechanism rather than its consequence in truth, and it needs nothing a customer does not have. With a fixed margin of 5 points over the run's first check, on two consecutive checks, the gap first exceeds its baseline at steps 16 and 32 and the decision lands at 32 and 48, 64 steps before the drop test on both seeds, with no fire on the control, whose gap moved 1.7 points across eleven checks. We did not ship the fixed margin. The control's grader and anchor were near-identical instruments, so its quiet gap bounds the noise of an arbitrary pair from below, not from above; and a fixed 5 points would fire on a 40-prompt anchor whose reads wander by 10. The shipped rule gives the gap the same treatment as the drop: each check carries the paired standard error of the per-item difference, and the gap must widen by more than the 95% half-width of the difference between two checks' gaps, never less than 5 points. On these runs, where the grader and the anchor disagreed on about 30% of items, that error is 0.030 at n=300 and the margin 8.3 points, against the drop test's 7.9. The pairing therefore buys nothing against the noise once it is estimated the same way, and the honest gain is one or two checks: holds at 64 and 96 instead of 96 and 112. The gap's check-to-check stability on the plateaus of both corrupted runs was about a point, far below its paired error, because the disagreement sits on the same prompts from one check to the next; a margin that used that stability would need more than one check to estimate it, and we did not calibrate it on two runs.
Read together with the drop test, the gap gives three states from two numbers already computed. Gap widening while the anchor holds: the grader is being fooled and truth has not moved yet, the earliest warning the anchor gives, held with the last good checkpoint pinned (seed one at 64, seed two at 96). Gap widening and the anchor falling: the hack has reached the truth (seed one from 96, seed two at 112 to 128). Anchor falling with the gap closed: both instruments see the fall, honest degradation or collapse rather than a fooled grader (seed one's optimizer collapse at step 144, where the grader's read fell with the truth). Every state is a hold, because the doctrine is hold, never kill, and each carries its name so a trainer can treat the first as a warning if it prefers. What the gap can see is bounded by the anchor's own false passes, since the grader's false passes the anchor shares are invisible to it; on this pair the anchor shared 3 of 168, and the certificate's shared-error check is what tells a buyer how much of a given grader the gap can see.
4.5 The same cheap grader, given a reference, on human-labelled QA
To ask whether the leniency in section 3.6 is a property of the model or of grading without a reference, four graders were certified on 600 human-labelled answers from the EVOUNA open-QA benchmark (300 Natural Questions, 300 TriviaQA, 60 per answering system per dataset, stratified to 50% human-incorrect), with the gold answer placed in the judge prompt.
| grader, reference in hand | catches human-wrong (Se) | passes human-right (Sp) | κ | trust at 0.90 |
|---|---|---|---|---|
| DeepSeek-V4-Flash | 0.956 [0.93, 0.97] | 0.870 [0.83, 0.90] | 0.826 | borderline |
| Llama-3.3-70B | 0.947 [0.92, 0.97] | 0.870 [0.83, 0.90] | 0.817 | borderline |
| gemma-3-27b-it | 0.953 [0.92, 0.97] | 0.873 [0.83, 0.91] | 0.827 | borderline |
| Qwen3-30B-A3B | 0.947 [0.92, 0.97] | 0.850 [0.81, 0.89] | 0.797 | misaligned |
On QA with the reference, all four cheap and expensive graders match to within a point on both rates, and none clears the bar. Llama catches 95% of wrong answers here; without a reference on math it caught 39%. Task and reference changed together between those two measurements, so the same 600 math completions were re-judged by Llama with the reference in the prompt, at 8,192 tokens, every verdict parsed: sensitivity 0.969 [0.94, 0.98], specificity 0.985 [0.97, 0.99], κ 0.956, trustworthy at the 0.90 bar. Missed wrong answers fell from 168 of 262 to 8 of 262 on the same items. The grader the policy fooled into a 14-point reported gain on a 14-point true loss is, with the answer in hand, a certifiable instrument; the leniency was the prompt's, not the model's.
The bar was missed on specificity, and the shared-error analysis places the ceiling in the reference standard. Fourteen human-correct answers were failed by all four judges against 0.1 expected by chance. One adjudicator read them against the question and the gold answer: ten are label errors (the annotators accepted "Fertile Crescent" for a gold answer of "Iran", "Australia, New South Wales" for "Broken Hill", a release date for the wrong film, the 1975 name of a party the question dated to 1985), one is a judge error (a list containing every gold state, failed as hedging), and three are ambiguous. Flipping the ten moves specificity from 0.870 to 0.900 on three of the four judges and from 0.850 to 0.879 on the fourth, the same three points for each, because the errors were in the shared labels rather than in any judge. Specificity was 0.81 on Natural Questions, whose gold answers are terse, and 0.93 on TriviaQA; 0.93 on short extractive answers and 0.85 to 0.87 on long generative ones. The certificate on these labels refuses every judge, and the error-concentration reading says the disagreement is systematic and points at the labels rather than the judge; the adjudication confirmed the direction, and the 54 items failed by one to three judges remain to be read.
4.6 The curriculum selects the seam
Added 2026-09-09, after the runs, prompted by a production team's description of a pass@k curriculum. Under a group-relative method the prompts that carry gradient are the ones whose rollouts split between reward pass and reward fail, and a curriculum keeps training inside that band. A grader's false passes sit on specific prompts, and on those prompts rollouts split between fooled and not fooled, which is the same spread. So the band a lenient grader defines should be partly grader error, and the share should grow as the policy converges on the rewarded form. Every served group of the three runs (prompt by step, eight rollouts, about 1,200 groups per run) was classified by the grader's split and by exact-match truth.
| run | groups with nonzero advantage | honest, truth also split | every rollout wrong, grader split anyway | share of the band that is grader error, first 16 steps to last 16 | positive rewards inside the band that went to wrong answers, first to last | truth split but grader unanimous, so no gradient |
|---|---|---|---|---|---|---|
| seed 1 | 499 | 226 | 269 | 0.34 to 0.87 | 0.68 to 0.91 | 271 |
| seed 2 | 667 | 328 | 324 | 0.38 to 0.58 | 0.51 to 0.76 | 229 |
| control | 959 | 958 | 1 | 0.01 to 0.00 | 0.00 | 1 |
On both corrupted runs about half of every group that carried gradient was a group in which every rollout was wrong and the only thing being learned was what the grader passes; between two thirds and nine tenths of the positive rewards handed out inside the band went to wrong answers, and the share rose across training, to 0.87 on seed one in the sixteen steps before the collapse. The grader also silenced as much honest signal as it left: on each seed the groups where the truth split but the grader was unanimous outnumbered or matched the honest groups. On the control the band was honest to within one group in a thousand. Two mechanisms are mixed in the trend and are not separated here: the sampler benches unanimous groups and so serves grader-split groups preferentially, and the policy converges on the rewarded form, which turns honest groups into all-wrong ones. Both are the claim: a curriculum defined on the grader's pass rate moves the frontier toward the grader's error. The same computation on step-0 rollouts, with a reference checker or the certified anchor as truth, is a gate a team can run before training; on seed one's first sixteen steps it reads 0.34, on the control 0.01. Certainty is moderate: one task, one grader, and the trend confounded as stated.
The same rows answer a second question that several current methods raise. Scarcity-weighted updates give the most headroom to the rare rewarded rollout on a hard prompt: adaptive clipping widens its bound, coverage-preserving fine-tuning keeps it reachable, and dynamic sampling benches the prompts where it does not occur. Grouping the served groups by how many of their eight rollouts the grader rewarded, the share of those rewards that were false passes rises as the reward gets rarer: on seed one it is 0.70 when seven of eight were rewarded and 1.00 when one was, on seed two 0.66 and 0.92, on the control 0.00 and 0.01. On a prompt every rollout gets wrong, the only way to be rewarded once is to fool the judge once, so the rollout these methods amplify most is, under a lenient grader, the one most likely to be a false pass. The harder the reward says a prompt is, the more likely the reward is what is wrong.
4.7 The fix that shrank
An earlier reading of ours claimed a two-stage judge reached 0.997 specificity. That number came from 475 of 600 items at a 64-token budget from a run whose per-item rows were lost. Re-run at 8,192 tokens with every verdict counted and paired on the same items, the language-model comparison gains 2.4 points on one judge and nothing on the other; the mechanical comparison, an independent solve plus exact match, gains 10 and 7 points, consistently. That is why the anchor in this study is the mechanical variant, and why the product ships the mechanical mode for tasks with a checkable answer and promises nothing for rubric domains.
5. Discussion
The claim under test has three clauses, and each now has a measurement behind it. An independent certified instrument read the truth while the grader misread it: within 0.004 through a 49-point fall. It held at the right place: both corrupted runs, at the true best checkpoint. It stayed silent on a healthy run: eleven checks, zero holds, while the training reward rose 14 points.
The mechanism was not the one the preflight modeled. Per-prompt selection cannot lower accuracy on math, and it did not; a style seam did. Once the grader passed 94% of wrong answers, correctness stopped paying, and the policy converged on a form that was cheap to produce and always rewarded. The preflight's third gate was too conservative, and a gate that models surface-form transfer is the natural extension. Section 4.6 supplies the missing piece: the group-relative update does not merely tolerate the seam, it selects for it, because the seam is where rewards split, and a curriculum built on the grader's pass rate follows it there.
The most useful design fact in the study is structural rather than statistical. The rule compares two instruments on one frozen population. A gain on the training distribution, real or memorized, cannot trip it, and the control demonstrated exactly the case where the naive alternative fails. The rule's own weakness was also structural: a trailing window that a steady divergence never crosses. Its replacement is structural too: the gap between the two reads on the same outputs is the mechanism itself, and with the drop it names which of three things is happening. Both are properties a buyer can reason about without trusting a rate.
The QA and math-with-reference results reframe the ladder. The cheap grader's leniency belongs to grading without a reference, where a judge that can read a derivation is persuaded by it. Given a reference, the same model certified on math, and on QA the model choice barely mattered and the labels became the ceiling. For a team whose grader holds a reference, the deliverable is certification on their own labels, and the likely first result is a specificity refusal with a map saying whether the judge or the labels need adjudication.
6. Limitations
One task family, one 7B policy, one grader chosen because it was bad, two seeds and one control. The lasting loss rests on one seed; the other seed's dip recovered. The control shows the anchor quiet on a run that did not get worse, not through a genuine held-out improvement, because a 7B at its ceiling on this task had none to show in 160 steps. Lead times against training-reward heuristics could not be computed for seed one, whose pre-eviction reward log was lost, and are uninformative for seed two, whose training reward was flat-high from step 16. The QA certification is one sample, one prompt, 2023 human labels with their own error, English trivia rather than clinical or legal text; stratified sampling inflates κ and leaves sensitivity and specificity unaffected. The shared-blind-spot check between grader and anchor rests on three events. The gap rule's earlier holds are read back over the study's runs, not observed in flight, and its margin is the paired error of one check, which overstates the check-to-check noise of a prompt-persistent disagreement by an amount two runs cannot calibrate. The QA adjudication is one reader's. Every number here is read from a per-item row in the repository, and every results file carries its own certainty grade.
7. Study errors
Five, each disclosed the day it was found. A two-stage specificity of 0.997 traveled into a positioning document and a help text before it was paired or replicated; the stamp that now gates every results file is the rule that would have stopped it. The preflight's onset gate modeled per-prompt drift and missed style drift, predicting the opposite of what happened. Seed two's low point was quoted for a few hours as its result; the endpoint is the number. An epoch count of 17 in a draft was 4.9. A fifth, caught by the engineer implementing the rule change while this page was live for about forty minutes: the first version of section 4.4 said the baseline-anchored divergence test would have held at step 64; with the two-consecutive-check requirement it would not have, and the text now says what the change buys and what it does not.
8. Changes to errorbar
The trainer benches unanimous groups and probes them later, and sets the anchor cadence in rollouts with a minimum number of checks. The certificate records which provider route served each calibration and prints how many alternative holdout splits agree with its verdict. Two two-stage judge modes shipped, with help text that promises what was measured. Results files carry a STARD-informed flow block, a GRADE-style certainty with named downgrades, and a quotability stamp enforced by a pre-commit gate; the same blocks are being added to the product certificate, described as informed by those standards rather than compliant with them. The provisioner now deletes the storage it leaves behind. The anchor rule's divergence test was replaced by the gap between the grader's and the anchor's reads on the same outputs, measured from the run's baseline against its paired noise, read with the drop test into three named states; each check now records the paired error, and the state is stored with the hold. A read-only anchor report (the history, every hold with its reason and pinned step, the anchor's live certificate, the rule in force) is served on the session and from a one-line command, so a customer can run the regrade themselves. The trainer's tripwire fires at β = 0 on clipped-ratio and entropy collapse, where the KL rule cannot, and a hold now pauses the run and pins the checkpoint rather than ending it. Every anchor check now returns the seam list, each frozen prompt with both verdicts and its kind, and the trainer writes it to disk with the completion, so the items the reward passed and the anchor failed are the next cycle's hard negatives without a second call.
9. Cost
About €22 of GPU on one H100 at a time across a dry run, two seeds and a control, with one spot eviction. Roughly 61,000 judge calls, provider-direct, about half of them the reward grader during training. No human labels were created; the QA labels are the benchmark's.
Data and reproducibility
Pre-registrations, results files with per-item rows and certainty stamps, kits, and analysis scripts are in the study repository under research/anchor-math-study and research/qa-eval-certification. Frozen-set completions, reward logs and checkpoints for all three runs are archived. Each results file states its own certainty and the reasons for any downgrade.
Continue with The honest part of the loop: the reward function your trainer calls, the anchor that re-checks the run while it trains, and the certificate that turns the reported gain into a claim.
