Towards AITowards AIToneBench

Thinking levels · OpenAI

Does GPT-5.4 write better when it thinks harder?

By a lot. We ran GPT-5.4 at its default and at xhigh, and xhigh lands 328 Elo higher with ranges far apart, strong evidence the jump is real. It costs 2.4 times as much and takes 3.8 minutes a script instead of 1.1.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptBoard rank
xhigh1888 ±66-85.0$0.1923.8 min#67
default (no effort flag)1559 ±43-82.0$0.0801.1 min#93

Which setting should you use?

If you're using GPT-5.4, don't leave it on the default. The biggest piece of the gap is length: at default it misses the target length badly, and xhigh fixes a good part of that while trimming some of the slop too. The voice is decent at both settings, so you're mostly paying for a script that fits. That said, I wouldn't start anything new on it. GPT-6 Sol (high) gets much closer to the target length, which is the very thing GPT-5.4 struggles with, and it ranks higher (#47 against #67) for $0.060 a script instead of $0.192.

Other OpenAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.