Submit your model to the bench.
An independent benchmark for code-generating language models, where the score is machine-checked — compiled and mathematically proven — not self-reported. You bring the key; we bring the gate. The suite is growing, and models measured early become the reference the rest are read against. The current table is at results/models.
What this is
Most model benchmarks are self-reported, or graded on a judgement call. Ours isn’t. We set your model a real software specification and put what it produces through a compiler and a formal proof of correctness — the same gate every piece of software we ship has to clear. The code either compiles and proves, or it does not. There is no partial credit and no talking your way to a good score. We record token cost alongside the result, so the reading is of capability and efficiency, not a single rank.
Why we ask you to bring the key
Our tests are large and getting larger, and running a model through them costs real inference. Rather than bankroll it — and, more to the point, rather than let anyone wonder whether the party paying for a run is the party scoring it — we ask each provider to supply a key for their own model. You fund your own inference; we stay neutral by construction. It is the same discipline that lets us publish a number anyone can trust.
The promise
- Every model runs through the identical harness. Rank cannot be bought.
- We never publish the harness or the prompts — so no model can be quietly tuned to the test. What is public is the result, never the mechanism.
- We publish honestly, including when we are wrong about a model. We have corrected, and apologised by name, for a benchmark error that wrongly marked two models down — we would sooner do that in the open than amend a table quietly. The rule cuts both ways: you cannot buy a good score here, and we will not let our own defect stand as your model’s failure.
Your key
Used only to run your model through the benchmark. Held outside version control, never redistributed, and yours to rotate the moment the run is done. If you would rather scope or rate-limit a key to the run, that is welcome.
To take part
→ tony.gair@thedarkfactory.co.uk
Tell us the model and the endpoint; we’ll tell you what the run involves before anything is spent. UK timezone; a human reads every message.
How the factory itself is built is described in our methodology.