Towards AITowards AIToneBench

Thinking levels · Anthropic

Does Claude Sonnet 5 write better when it thinks harder?

Clearly, but you wait for it. Max takes Claude Sonnet 5 from #73 up to #48 on the board, 221 Elo above its default, a real jump that still leaves it well short of the top. Each script then takes about 11.4 minutes instead of 1.3, and costs 8.0 times as much.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptBoard rank
max2034 ±48-86.4$0.73311.4 min#48
default (no effort flag)1814 ±41-84.5$0.0911.3 min#73

Which setting should you use?

Max is the better setting here, and the ranges sit far enough apart to call that strong evidence. I just don't think either one is the right call anymore. Claude Sonnet 5.5 (high) ranks above this model at max, costs $0.077 a script against $0.733, and finishes in 0.7 minutes instead of 11.4. If you're on this model today, upgrading to the newer Sonnet does more for your scripts than turning up the effort.

Other Anthropic ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Anthropic model side by side.