Towards AITowards AIToneBench

Thinking levels · Moonshot (Kimi)

Does Kimi K3 write better when it thinks harder?

Yes, but the labels point the wrong way. Plain Kimi K3, with no effort flag, spent far more reasoning tokens than Kimi K3 (thinking) and scored 140 Elo higher, with ranges that don't overlap. That extra thinking isn't free: 2.7 times the cost per script, and 3.9 minutes of waiting instead of 1.2.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
high2068 ±40-86.9$0.0951.2 min1,268#43
default (no effort flag)2208 ±37-88.1$0.2603.9 min12,239#21

Which setting should you use?

I'd run Kimi K3 with no flag for anything I'm going to publish. The board gives strong evidence that it writes the better script, and at $0.260 it's still a cheap model to use. The thinking version makes sense when you need lots of drafts fast: it sits at #43 on the board for $0.095 and comes back in 1.2 minutes. Just don't pick it because the name says thinking. On this model, the default run is the one that actually thinks harder.

Other Moonshot (Kimi) ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Moonshot (Kimi) model side by side.