Moonshot's newest model writes well: Kimi K3 sits at #21 on the whole board, and like every Kimi here it's open weights, so you can host it yourself. Then there's a cliff. The older Kimi K2 rows trail far behind, and from the top of the family to the bottom the gap reaches 920 Elo. So for most people, the real choice is between the two Kimi K3 settings.
Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.
| Board rank | Configuration | Elo | Overall | Tone | Craft | Substance | Hook | YouTube | Flow | Slop | Length | Cues | Cost / script | Time / script |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| #21 | Kimi K3 · open | 2208 ±37 | 88.1 | 88.2 | 89.0 | 86.9 | 88.6 | 87.7 | 87.2 | 90.1 | 90.1 | 84.9 | $0.260 | 3.9 min |
| #43 | Kimi K3 (thinking) · open | 2068 ±40 | 86.9 | 87.2 | 88.4 | 86.6 | 87.7 | 86.5 | 86.3 | 88.6 | 86.3 | 83.3 | $0.095 | 1.2 min |
| #91 | Kimi K2.6 · open | 1575 ±41 | 82.0 | 82.5 | 83.1 | 81.4 | 85.8 | 81.3 | 79.5 | 84.1 | 84.8 | 71.9 | $0.078 | 4.4 min |
| #92 | Kimi K2.6 (thinking) · open | 1567 ±37 | 81.8 | 82.3 | 82.8 | 81.1 | 85.5 | 80.0 | 79.6 | 84.5 | 85.1 | 73.1 | $0.083 | 4.3 min |
| #114 | Kimi K2 (0905) · open | 1323 ±45 | 78.1 | 79.1 | 81.9 | 78.5 | 83.5 | 73.0 | 77.1 | 83.1 | 69.7 | 75.1 | $0.012 | 0.8 min |
| #117 | Kimi K2 Thinking · open | 1288 ±46 | 78.0 | 78.7 | 80.6 | 77.7 | 84.2 | 75.5 | 76.6 | 80.9 | 71.9 | 72.6 | $0.026 | 3.0 min |
Kimi K3 is the one I'd publish with, at $0.260 per script. If you write a lot, Kimi K3 (thinking) is the value pick. It's 140 Elo behind, and here's the odd part: plain Kimi K3 costs 2.7 times as much and takes 3.9 minutes per script against 1.2 minutes, so thinking mode is the cheaper, faster option this time.
Funny thing about this family: none of the thinking variants clearly beats its plain sibling, so don't flip the switch by reflex.
Kimi K2 (0905) is the budget pick at $0.012, but I'd skip the older rows anyway: Kimi K2.6 already costs close to what the thinking Kimi K3 charges, for a much weaker script.
Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.
Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.
← All comparisons