errorbar finds behaviors frontier models are missing.

We confirm the gaps with controlled counterfactuals, and build software-verifiable RL environments to train them.

Where we start

We start in finance and accounting, where professional work can be realistic while correctness remains exactly verifiable.

The loop

How a gap becomes an environment, and where the next one comes from

  1. 01

    Test

    We test frontier models on real professional work.

  2. 02

    Confirm

    We use controlled counterfactuals to distinguish systematic behavioral gaps from one-off failures.

  3. 03

    Build

    We turn confirmed gaps into software-verifiable training environments where correctness is computed directly, not judged by humans or models.

  4. 04

    Search again

    Each confirmed gap steers the next search. Back to step one.

First confirmed gap

Verified Completion

Models can often do the mechanical work correctly, but fail to recognize when the evidence is insufficient to sign off.

What the work requires

The evidence contradicts a figure
Correct it.
The evidence supports it
Sign off.
The evidence is missing
Do not sign off.Models often sign off anyway.

Get in touch