Towards AITowards AIToneBench

Thinking levels · OpenAI

Does GPT-5.4 mini write better when it thinks harder?

Slightly, and I wouldn't call that gap settled. Going from the plain GPT-5.4 mini to high adds 76 Elo, for 1.5 times the cost and 2.4 minutes a script instead of 1.5.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptBoard rank
high1523 ±43-81.4$0.0532.4 min#97
default (no effort flag)1447 ±41-80.4$0.0361.5 min#107

The 95% Elo ranges overlap: GPT-5.4 mini runs from 1401 to 1483, and GPT-5.4 mini (high) from 1479 to 1566.

Which setting should you use?

On mini, use high anyway: its scripts land closer to the target length, and the extra cost is small. But I'd question why you're on mini at all. GPT-6 Sol (default) costs about the same at $0.054 a script and ranks far higher, #55 against #97. And if cost is the real constraint, GPT-6 Luna (max) writes better than either mini setting for $0.0041. On scripts, mini gets beaten from both sides.

Other OpenAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.