Towards AITowards AIToneBench

Thinking levels · Moonshot (Kimi)

Does Kimi K2.6 write better when it thinks harder?

Not really, and the flag barely changes anything here. The two settings land 8 Elo apart with overlapping ranges, cost about the same per script, and both take around 4.4 minutes. Both runs already spend a lot of reasoning tokens, so there just isn't much left for the setting to move.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
high1567 ±37-81.8$0.0834.3 min15,604#92
default (no effort flag)1575 ±41-82.0$0.0784.4 min14,305#91

The 95% Elo ranges overlap: Kimi K2.6 runs from 1533 to 1616, and Kimi K2.6 (thinking) from 1528 to 1603.

Which setting should you use?

I'd leave it on default and stop thinking about it. The bigger decision is the model itself. Kimi K2.6 is #91, while Kimi K3 (thinking) sits far above it at #43 for $0.095 per script, and it's faster too. So if you're on Kimi K2.6 today, I'd switch models before I'd tune this one.

Other Moonshot (Kimi) ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Moonshot (Kimi) model side by side.