Benchmark

Submit your model to the bench.

An independent benchmark for code-generating language models, where the score is machine-checked — compiled and mathematically proven — not self-reported. You bring the key; we bring the gate. The suite is growing, and models measured early become the reference the rest are read against. The current table is at results/models.

What this is

Most model benchmarks are self-reported, or graded on a judgement call. Ours isn’t. We set your model a real software specification and put what it produces through a compiler and a formal proof of correctness — the same gate every piece of software we ship has to clear. The code either compiles and proves, or it does not. There is no partial credit and no talking your way to a good score. We record token cost alongside the result, so the reading is of capability and efficiency, not a single rank.

Why we ask you to bring the key

Our tests are large and getting larger, and running a model through them costs real inference. Rather than bankroll it — and, more to the point, rather than let anyone wonder whether the party paying for a run is the party scoring it — we ask each provider to supply a key for their own model. You fund your own inference; we stay neutral by construction. It is the same discipline that lets us publish a number anyone can trust.

The promise

Your key

Used only to run your model through the benchmark. Held outside version control, never redistributed, and yours to rotate the moment the run is done. If you would rather scope or rate-limit a key to the run, that is welcome.

To take part

tony.gair@thedarkfactory.co.uk

Tell us the model and the endpoint; we’ll tell you what the run involves before anything is spent. UK timezone; a human reads every message.

How the factory itself is built is described in our methodology.