Thinking levels · Anthropic
A little at every step, but only the last one is a clear gain. From low to max, Claude Opus 5 gains 182 Elo, so in a straight script comparison max comes out ahead of low about 74% of the time. What's unusual is how little that costs here: max is 1.9 times the price of low, and a script still comes back in about 2.2 minutes.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Board rank |
|---|---|---|---|---|---|---|
| low | 2160 ±39 | - | 87.6 | $0.151 | 1.0 min | #29 |
| medium | 2176 ±39 | +16 | 87.8 | $0.154 | 1.0 min | #25 |
| high | 2229 ±35 | +54 | 88.2 | $0.171 | 1.2 min | #17 |
| xhigh | 2253 ±38 | +23 | 88.4 | $0.203 | 1.4 min | #14 |
| max | 2342 ±41 | +89 | 89.0 | $0.293 | 2.2 min | #8 |
| default (no effort flag) | 2242 ±35 | - | 88.3 | $0.176 | 1.2 min | #16 |
One disclosure: Claude Opus 5 itself also holds a seat on the judge panel. The panel spans three LLM families plus Jev, so no family scores itself alone, and the methodology explains how that works.
Run it at max. The lower steps overlap with each other, but max's range sits clear of xhigh's, and you're paying $0.293 a script for #8 on the board. That's a cheap top setting. If you're picking a model today though, Claude Opus 5.5 (high) sits inside the same Elo range, costs about the same and comes back a bit faster, so that's my pick for a fresh start.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Anthropic model side by side.