Towards AITowards AIToneBench

Thinking levels · OpenAI

Does GPT-5.6 Luna write better when it thinks harder?

A lot, actually, and almost all of it in the last step. GPT-5.6 Luna climbs 281 Elo from no reasoning to ultra, enough that the top setting would win about 83% of script comparisons against the bottom one. Ultra still costs $0.023 a script, so what you really pay is the wait: 5.5 minutes, up from 1.0 with no reasoning.

X axis

Every setting, step by step

SettingEloStepOverallCost / scriptTime / scriptBoard rank
no reasoning1513 ±44-81.4$0.00611.0 min#98
low1447 ±29-6580.7$0.00631.1 min#105
high1534 ±46+8781.7$0.00711.3 min#95
xhigh1600 ±43+6682.4$0.00921.9 min#89
ultra1793 ±40+19484.3$0.0235.5 min#74
default (no effort flag)1477 ±37-81.0$0.00651.1 min#102

Which setting should you use?

Ultra, for me, unless you need scripts fast. The length scores explain a lot of the jump: below ultra, GPT-5.6 Luna keeps missing the script's target length, and at ultra it mostly lands it. The no reasoning and ultra ranges don't overlap, so that's strong evidence the gain is real. If the wait is a problem, xhigh comes back in 1.9 minutes, but it gives up most of the climb. And keep expectations in check: even at its best, GPT-5.6 Luna ranks #74 on the board, so think of it as a budget model on its best setting.

Other OpenAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.