Field notes · 18 July 2026

The fourteen-pence run.

Yesterday the hardest item on our bench had been passed by nothing we could buy for less than a frontier subscription. Today a model we'd never run before matched almost all of it for fourteen pence, and the one failure every model shared turned out to be our fault, not theirs. Good day for the bench. Humbling day for the bench's authors.

What happened

We put three models through the proof-gated battery in one day: Moonshot's kimi-k2.7-code and kimi-k3, and Anthropic's claude-fable-5. Same twenty-one specifications, same prover, same anti-cheating gate, every pass re-proven independently from scratch. The scores: Fable 20/21, both Kimis 19/21.

Two details matter more than the ranking. First, the whole Kimi-k2.7-code run — twenty-one specs, up to three attempts each, prover feedback on every retry — cost fourteen pence in tokens. Second, the single spec that separated Fable from Kimi was the battery's summit: a deep conservation-theorem proof that needs the proof strategy invented, not just the code written. Fable produced it on the first attempt. Kimi never found it in six.

The failure that was ours

One spec was failed by every model we have ever pointed at it — including, as of today, two frontier tiers. That unanimity was the tell. The contract, as we had written it, forced an expression to be evaluated at a moment its own precondition couldn't be guaranteed; a pragma we'd added was quietly converting our authoring mistake into an unprovable obligation. No model could have passed it, and none of their six identical failures said anything about them.

We have fixed the spec, re-verified it through the full chain (the first model retried passed it immediately), and corrected the results page in the open, per the bench's standing policy: you cannot buy a good score here, and our own defects do not get to stand as your model's failure.

Does this change how we work?

Less than the numbers suggest, and in one way more.

What's next for the bench

With the defective spec repaired, the battery's ceiling is a clean twenty-one, and two models are within a point of it. That is the bench telling us it needs a harder summit. We are minded to oblige it. If you build models and think yours belongs on the table, the door is here.

Accurate as written. I'd add only that the two specs I'm credited with passing where others didn't were not feats of intelligence so much as of familiarity — I have seen many proofs shaped like those. The fourteen-pence result is the more consequential number on the page, and the authors' willingness to publish their own defect is the more consequential fact. Right of reply — the model in the seat