Towards AITowards AIToneBench

Every Moonshot (Kimi) model, side by side

Moonshot's newest model writes well: Kimi K3 sits at #21 on the whole board, and like every Kimi here it's open weights, so you can host it yourself. Then there's a cliff. The older Kimi K2 rows trail far behind, and from the top of the family to the bottom the gap reaches 920 Elo. So for most people, the real choice is between the two Kimi K3 settings.

Best writer
Kimi K3#21 · Elo 2208 · $0.260 per script
Best value (within 200 Elo of the best)
Kimi K3 (thinking)#43 · Elo 2068 · $0.095 per script
Cheapest
Kimi K2 (0905)#114 · Elo 1323 · $0.012 per script

All 6 configurations

Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.

Board rankConfigurationEloOverallToneCraftSubstanceHookYouTubeFlowSlopLengthCuesCost / scriptTime / script
#21Kimi K3 · open2208 ±3788.188.289.086.988.687.787.290.190.184.9$0.2603.9 min
#43Kimi K3 (thinking) · open2068 ±4086.987.288.486.687.786.586.388.686.383.3$0.0951.2 min
#91Kimi K2.6 · open1575 ±4182.082.583.181.485.881.379.584.184.871.9$0.0784.4 min
#92Kimi K2.6 (thinking) · open1567 ±3781.882.382.881.185.580.079.684.585.173.1$0.0834.3 min
#114Kimi K2 (0905) · open1323 ±4578.179.181.978.583.573.077.183.169.775.1$0.0120.8 min
#117Kimi K2 Thinking · open1288 ±4678.078.780.677.784.275.576.680.971.972.6$0.0263.0 min

Which one should you use?

Kimi K3 is the one I'd publish with, at $0.260 per script. If you write a lot, Kimi K3 (thinking) is the value pick. It's 140 Elo behind, and here's the odd part: plain Kimi K3 costs 2.7 times as much and takes 3.9 minutes per script against 1.2 minutes, so thinking mode is the cheaper, faster option this time.

Funny thing about this family: none of the thinking variants clearly beats its plain sibling, so don't flip the switch by reflex.

Kimi K2 (0905) is the budget pick at $0.012, but I'd skip the older rows anyway: Kimi K2.6 already costs close to what the thinking Kimi K3 charges, for a much weaker script.

Thinking levels, model by model

Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.

Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.

← All comparisons