Towards AITowards AIToneBench

Thinking levels · Anthropic

Does Claude Sonnet 5.5 write better when it thinks harder?

Massively. Claude Sonnet 5.5 goes from #70 on the board at low effort to #3 at max, a gap wide enough that max should beat low on about 98% of scripts. You pay for that jump, though: 23.7 times the price per script, and 12.9 minutes of waiting instead of 0.4.

X axis

Every setting, step by step

SettingEloStepOverallCost / scriptTime / scriptBoard rank
low1872 ±50-85.0$0.0560.4 min#70
medium1957 ±42+8585.9$0.0590.5 min#54
high2136 ±33+17987.5$0.0770.7 min#33
xhigh2303 ±46+16788.7$0.2322.2 min#9
max2512 ±41+20990.1$1.3212.9 min#3
default (no effort flag)1918 ±39-85.6$0.0590.5 min#63

Which setting should you use?

The setting I'd avoid is the one you get by default. It sits with low and medium, which the board can't separate, so passing no effort flag leaves most of this model on the table. Claude Sonnet 5.5 (high) is my pick: #33 for $0.077, only pennies above the default. With a bit more budget, xhigh is another clear step up, at #9.

Max is a top-of-board writer, but I wouldn't run it. Claude Opus 5.5 (xhigh) lands in the same range, costs less, and finishes in 4.4 minutes instead of 12.9.

Other Anthropic ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Anthropic model side by side.