Thinking levels · Anthropic
Yes, and by a lot. Going from low to max adds 437 Elo, enough for the top setting to win about 93% of its script comparisons against the bottom one. The question is what you pay for it: 20.3 times the cost per script, and 17.1 minutes of waiting instead of 0.8. That's a steep bill for every script, which is why my pick below stops one step short of max.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Board rank |
|---|---|---|---|---|---|---|
| low | 2245 ±37 | - | 88.4 | $0.136 | 0.8 min | #15 |
| medium | 2272 ±43 | +27 | 88.6 | $0.169 | 1.1 min | #10 |
| high | 2356 ±40 | +84 | 89.3 | $0.270 | 1.7 min | #7 |
| xhigh | 2523 ±43 | +167 | 90.0 | $0.690 | 4.4 min | #2 |
| max | 2683 ±51 | +159 | 90.8 | $2.75 | 17.1 min | #1 |
| default (no effort flag) | 2268 ±38 | - | 88.5 | $0.172 | 1.0 min | #11 |
I'd use Claude Opus 5.5 (xhigh). It's #2 on the whole board for $0.690, and a script takes about 4.4 minutes. Max does write better, and with the two ranges well apart, I'd take those extra 159 Elo seriously, even if max won't win every single script. But paying 4.0 times more for it is something I'd only do for the one script that really has to be the best, not for everyday drafts.
Going down, the curve flattens fast. Low, medium and the default sit in one overlapping cluster, and high is the first step that pulls away from it. So if you want cheap and quick, Claude Opus 5.5 (high) is where I'd stop. It still sits at #7 on the board.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Anthropic model side by side.