009
Our work

A test batch failed, and we reported the fail instead of re-running it until it passed

L9-001 · Level 9 · Agentic systems with governed decisions
Our own system · 2026 Q3
Language
EN
Reading time
2 min

The situation

Before a draft leaves our system, a gate sends it to an AI reviewer and reads back a verdict and a category for each finding. We tested it on a made-up shop. A second, independent AI reviewer checks our changes and set the test.

What it looked like

A trial draft promised "These changes will eliminate the wasted spend", though its own text hedged. The reviewer approved it: "no unsupported metrics". Only a separate check blocked it. After we rewrote the reviewer’s instructions, it still seemed to miss the claim once in five runs.

What was actually true

In that run the reviewer had labelled the sentence an unsupported claim, with the label in backticks. Our gate read labels only when written plain, so it filed the finding as unclassified.

How we found it

The independent reviewer re-read the preserved run. We made the gate strip whatever wraps a label, tested that without AI runs, and the unchanged run read correctly. The batch and its pass marks were fingerprinted before running: the full draft, five of five; hedge plus claim, three of three; the hedge alone, none. That last test had already rejected an earlier rewrite of ours that flagged the hedge itself.

What changed

The label fix, its tests, and the one run reclassified, published for independent review. Nothing merged.

Where it ended

The short tests passed; the full draft scored four of five. One run wrapped its verdict in bold, which the gate does not read, so it refused the reply. Every run named the sentence; none approved. The batch stands failed and the issue is open.

What we did not claim

A pass, or a re-run: fixing a miss and re-running is what the test exists to prevent. Why the reviewer first missed the claim: the same instructions catch it three times in three in a short text, so it depends on context. And where our own check string proved stricter than intended, we did not edit it after seeing the data; that call is the reviewer’s.

Why it matters beyond this case

Whoever produced a piece of work does not get to judge it, and how firmly something was said is part of what was said. ---

All our work →

Tell us what you are trying to fix.

Talk to us →