Thinking levels · Anthropic
Quite a bit. Max takes Claude Opus 4.8 from #46 to #23, 146 Elo above the default, which is the only other setting we ran. That jump costs 4.2 times more per script, and you wait about 5.2 minutes instead of 1.1.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Board rank |
|---|---|---|---|---|---|---|
| max | 2187 ±44 | - | 87.8 | $0.678 | 5.2 min | #23 |
| default (no effort flag) | 2042 ±37 | - | 86.6 | $0.160 | 1.1 min | #46 |
If you're stuck on Claude Opus 4.8, run it at max. The default and max ranges never touch, which is strong evidence the gain holds up. For anything new, though, I'd skip this ladder entirely: Claude Opus 5.5 (adaptive default), with no effort flag at all, already ranks above Claude Opus 4.8 (max effort), and the older model at max costs 3.9 times as much per script. The newer Opus is just a better deal for scripts.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Anthropic model side by side.