Towards AITowards AIToneBench

Thinking levels · Anthropic

Does Claude Fable 5 write better when it thinks harder?

Only once you get to xhigh. Below that, the settings barely move: low, medium, high and the default all land in the same range, and high even comes out a touch under medium. Then xhigh jumps to #6. From low to max, that adds up to 189 Elo, and max beats low in about 75% of script comparisons.

X axis

Every setting, step by step

SettingEloStepOverallCost / scriptTime / scriptBoard rank
low2207 ±38-88.0$0.3050.9 min#22
medium2221 ±37+1588.1$0.3110.9 min#20
high2187 ±38-3487.9$0.3191.0 min#24
xhigh2360 ±43+17289.3$1.023.2 min#6
max2396 ±43+3689.5$2.529.1 min#5
default (no effort flag)2222 ±34-88.1$0.5572.0 min#19

Which setting should you use?

I'd run it at xhigh. Max adds only 36 Elo on top, and the two ranges overlap, so the board can't separate them with confidence. And for that sliver you'd pay 2.5 times as much and wait 9.1 minutes instead of 3.2.

At the bottom, low is the only setting worth considering. The default costs 1.8 times as much per script without the board being able to tell it apart from low, and paying for high doesn't buy anything either. So either go straight to xhigh or stay at low.

Other Anthropic ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Anthropic model side by side.