Thinking levels · Anthropic
Only once you get to xhigh. Below that, the settings barely move: low, medium, high and the default all land in the same range, and high even comes out a touch under medium. Then xhigh jumps to #6. From low to max, that adds up to 189 Elo, and max beats low in about 75% of script comparisons.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Board rank |
|---|---|---|---|---|---|---|
| low | 2207 ±38 | - | 88.0 | $0.305 | 0.9 min | #22 |
| medium | 2221 ±37 | +15 | 88.1 | $0.311 | 0.9 min | #20 |
| high | 2187 ±38 | -34 | 87.9 | $0.319 | 1.0 min | #24 |
| xhigh | 2360 ±43 | +172 | 89.3 | $1.02 | 3.2 min | #6 |
| max | 2396 ±43 | +36 | 89.5 | $2.52 | 9.1 min | #5 |
| default (no effort flag) | 2222 ±34 | - | 88.1 | $0.557 | 2.0 min | #19 |
I'd run it at xhigh. Max adds only 36 Elo on top, and the two ranges overlap, so the board can't separate them with confidence. And for that sliver you'd pay 2.5 times as much and wait 9.1 minutes instead of 3.2.
At the bottom, low is the only setting worth considering. The default costs 1.8 times as much per script without the board being able to tell it apart from low, and paying for high doesn't buy anything either. So either go straight to xhigh or stay at low.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Anthropic model side by side.