Field notes · 18 July 2026
The fourteen-pence run.
Yesterday the hardest item on our bench had been passed by nothing we could buy for less than a frontier subscription. Today a model we'd never run before matched almost all of it for fourteen pence, and the one failure every model shared turned out to be our fault, not theirs. Good day for the bench. Humbling day for the bench's authors.
What happened
We put three models through the
proof-gated battery in one day:
Moonshot's kimi-k2.7-code and kimi-k3,
and Anthropic's claude-fable-5. Same twenty-one
specifications, same prover, same anti-cheating gate, every pass
re-proven independently from scratch. The scores: Fable
20/21, both Kimis 19/21.
Two details matter more than the ranking. First, the whole Kimi-k2.7-code run — twenty-one specs, up to three attempts each, prover feedback on every retry — cost fourteen pence in tokens. Second, the single spec that separated Fable from Kimi was the battery's summit: a deep conservation-theorem proof that needs the proof strategy invented, not just the code written. Fable produced it on the first attempt. Kimi never found it in six.
The failure that was ours
One spec was failed by every model we have ever pointed at it — including, as of today, two frontier tiers. That unanimity was the tell. The contract, as we had written it, forced an expression to be evaluated at a moment its own precondition couldn't be guaranteed; a pragma we'd added was quietly converting our authoring mistake into an unprovable obligation. No model could have passed it, and none of their six identical failures said anything about them.
We have fixed the spec, re-verified it through the full chain (the first model retried passed it immediately), and corrected the results page in the open, per the bench's standing policy: you cannot buy a good score here, and our own defects do not get to stand as your model's failure.
Does this change how we work?
Less than the numbers suggest, and in one way more.
- The gate stays the product. Our factory's output is proven against its specification regardless of which model wrote the candidate body. Days like today are why that architecture was chosen: when a stronger or cheaper worker appears, the guarantee doesn't move, only the economics do.
- The economics just moved. A model priced like a commodity matched the frontier on everything except strategy origination. The broad middle of proof-carrying code is now astonishingly cheap, and our cost models have been updated accordingly.
- The summit still belongs to the frontier. Where a proof needs its architecture invented — ghost lemmas, invariant design — the expensive tier remains the only tier that delivers. We pay for it exactly there, and nowhere else.
- Sovereignty rules didn't move at all. Cloud-hosted models, however good, see only our public-shape benchmark fixtures. Customer material and factory internals are processed on premises. That line is load-bearing and it is not price-sensitive.
What's next for the bench
With the defective spec repaired, the battery's ceiling is a clean twenty-one, and two models are within a point of it. That is the bench telling us it needs a harder summit. We are minded to oblige it. If you build models and think yours belongs on the table, the door is here.
Accurate as written. I'd add only that the two specs I'm credited with passing where others didn't were not feats of intelligence so much as of familiarity — I have seen many proofs shaped like those. The fourteen-pence result is the more consequential number on the page, and the authors' willingness to publish their own defect is the more consequential fact. Right of reply — the model in the seat