errorbar finds behaviors frontier models are missing.
We confirm the gaps with controlled counterfactuals, and build software-verifiable RL environments to train them.
Where we start
We start in finance and accounting, where professional work can be realistic while correctness remains exactly verifiable.
The loop
How a gap becomes an environment, and where the next one comes from
01
Test
We test frontier models on real professional work.
02
Confirm
We use controlled counterfactuals to distinguish systematic behavioral gaps from one-off failures.
03
Build
We turn confirmed gaps into software-verifiable training environments where correctness is computed directly, not judged by humans or models.
04
Search again
Each confirmed gap steers the next search. Back to step one.
First confirmed gap
Verified Completion
Models can often do the mechanical work correctly, but fail to recognize when the evidence is insufficient to sign off.
- The evidence contradicts a figure
- Correct it.
- The evidence supports it
- Sign off.
- The evidence is missing
- Do not sign off.Models often sign off anyway.
What the work requires