AI Tools

Which open-weight AI model to use for coding

Loading the index…

What this index cannot score

The honest shape of the evidence, because it is the most useful thing on this page: licence terms are knowable for every model here, and coding performance is knowable for almost none of them.

  • Licence terms: complete coverage. Every scored cell was read from the model’s own licence document, not a model card, announcement or aggregator.
  • Measured coding performance: one of the three in-scope models. Only DeepSeek V4 Flash has been put through a controlled Kingy run. The other two are marked Not yet verified, which means we have not measured them — not that they performed badly.
  • Published per-token pricing: one of five. The rest publish prose rather than a figure, or are meant to be self-hosted, where a per-token price does not apply.
  • Scope itself had to be corrected once already. Leanstral-1.5 was ranked third here until 8 August, on the same assumption the coverage it appears in makes. Reading the vendor’s own model card is what caught it, and the change is logged above rather than quietly applied.
  • That asymmetry is why the score is what it is. Ranking five models on a capability number we hold for one of them would be a fabrication. Ranking them on terms we verified for all five is not, and the terms turn out to be decisive under this profile anyway.

How to read this

  • It is one answer, not a shortlist. The full ranking is shown so the runner-up and the distance to it are visible, but the pick is the top row under the stated conditions.
  • The score is computed, the verdict is written. Every score reconstructs from the published rubric and the licence documents; verdicts are editorial and say what the number does not.
  • Scope gates before score. A model built for something other than repository coding ranks last whatever its licence scores, and its row says why. Leanstral-1.5 is the case in point: a legitimate 10.0 on terms, and a Lean 4 theorem prover. A licence score is not a recommendation on its own — that is the whole reason this column is named what it is.
  • Two in-scope models tie at the top. That is a real result, not a rounding artefact — on deployment latitude DeepSeek V4 Flash and GLM-5.2 are equivalent. The tie is broken by which one we have actually measured.
  • Cost is list price per million tokens where a first-party figure was verifiable, with the date it was checked. It is not a blended or negotiated rate, and not what self-hosting costs you.
  • Open weights are not one category. Two models with the same score can still carry very different obligations; the licence index is the full reading.

The source of truth is a dataset published alongside this page and read by your browser on every load, so what you see is its current state rather than a copy taken when the page was last edited. Rows still being worked on are held separately and are not read by this page.