Towards AITowards AIToneBench

Thinking levels · Anthropic

Does Claude Fable 5.1 write better when it thinks harder?

At max, yes, clearly. We only ran Claude Fable 5.1 at its default and at max, and max climbs from #13 to #4 on the board, 228 Elo higher. The bill is the hard part: $3.15 and about 10.4 minutes per script, 4.9 times the price of the default.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptBoard rank
max2481 ±44-90.0$3.1510.4 min#4
default (no effort flag)2253 ±42-88.4$0.6482.1 min#13

Which setting should you use?

If you're already on this model, max is the setting that earns its place, and its range sits far enough from the default's that I wouldn't put the jump down to luck. But I wouldn't start here today. Claude Opus 5.5 (max effort) ranks higher and costs less per script. And Claude Opus 5.5 with no effort flag overlaps with Fable's default, which costs 3.8 times as much. So the newer Opus wins at max, can't be told apart at the default, and costs less both times.

Other Anthropic ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Anthropic model side by side.