Thinking levels · Moonshot (Kimi)
Yes, but the labels point the wrong way. Plain Kimi K3, with no effort flag, spent far more reasoning tokens than Kimi K3 (thinking) and scored 140 Elo higher, with ranges that don't overlap. That extra thinking isn't free: 2.7 times the cost per script, and 3.9 minutes of waiting instead of 1.2.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Reasoning tokens | Board rank |
|---|---|---|---|---|---|---|---|
| high | 2068 ±40 | - | 86.9 | $0.095 | 1.2 min | 1,268 | #43 |
| default (no effort flag) | 2208 ±37 | - | 88.1 | $0.260 | 3.9 min | 12,239 | #21 |
I'd run Kimi K3 with no flag for anything I'm going to publish. The board gives strong evidence that it writes the better script, and at $0.260 it's still a cheap model to use. The thinking version makes sense when you need lots of drafts fast: it sits at #43 on the board for $0.095 and comes back in 1.2 minutes. Just don't pick it because the name says thinking. On this model, the default run is the one that actually thinks harder.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Moonshot (Kimi) model side by side.