Towards AITowards AIToneBench

Thinking levels · OpenAI

Does GPT-5.6 Terra write better when it thinks harder?

On GPT-5.6 Terra, the effort setting changes your bill a lot more than your script. Going from no reasoning to ultra moves it 16 Elo, and the best score actually comes from xhigh, one step below the top. Ultra, meanwhile, runs $0.159 a script and keeps you waiting 3.8 minutes for each one.

X axis

Every setting, step by step

SettingEloStepOverallCost / scriptTime / scriptBoard rank
no reasoning1943 ±42-85.8$0.0551.6 min#58
low1912 ±42-3185.4$0.0591.0 min#64
high1939 ±44+2785.7$0.0591.0 min#59
xhigh1976 ±38+3786.1$0.0661.5 min#51
ultra1959 ±40-1785.9$0.1593.8 min#53
default (no effort flag)1962 ±39-86.0$0.0581.0 min#52

The 95% Elo ranges overlap: GPT-5.6 Terra (none) runs from 1903 to 1986, and GPT-5.6 Terra (ultra) from 1920 to 2001.

Which setting should you use?

Keep the default. It sits 15 Elo under xhigh, close enough to be noise, and costs $0.058 a script. If you want the best-scoring setting anyway, xhigh adds very little to the price, so go for it. Ultra is the one I'd skip: 2.7 times the default price, and it doesn't score any better than xhigh. The ranges overlap all along this ladder, so on Terra I'd treat the effort knob as a cost knob and leave it alone.

Other OpenAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.