Towards AITowards AIToneBench

Thinking levels · OpenAI

Does GPT-5.5 write better when it thinks harder?

It does, and on GPT-5.5 you barely pay for it. Going from minimal to xhigh adds 120 Elo, while the price per script hardly moves, $0.172 at the bottom and $0.178 at the top, and the wait stays about the same too.

X axis

Every setting, step by step

SettingEloStepOverallCost / scriptTime / scriptBoard rank
minimal1492 ±39-81.2$0.1721.3 min#101
high1621 ±52+12982.7$0.1781.5 min#84
xhigh1612 ±41-982.6$0.1781.4 min#86
default (no effort flag)1550 ±41-81.9$0.1711.4 min#94

Which setting should you use?

I'd run it on high. It's the best-scoring setting here, and xhigh lands close enough that the difference sits inside the noise, so there's no reason to go further. Neither range overlaps the minimal one, so that's strong evidence the gain over minimal is real.

What holds GPT-5.5 back at every setting is length: its voice and writing scores are solid, but its scripts keep missing the target length, and more thinking only fixes part of that. Honestly, at $0.178 a script, I'd only keep it if it's already wired into your setup, because GPT-5.6 Sol (default) scores higher for $0.106.

Other OpenAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.