Towards AITowards AIToneBench

Thinking levels · OpenAI

Does GPT-6 Sol write better when it thinks harder?

It does. The max setting beats high by 77 Elo, so it would win about 61% of their script comparisons. You pay 1.9 times as much, but that still only comes to $0.112 a script, and the wait grows from 1.2 minutes to 2.9.

X axis

Every setting, step by step

SettingEloStepOverallCost / scriptTime / scriptBoard rank
high2037 ±37-86.6$0.0601.2 min#47
max2114 ±38+7787.2$0.1122.9 min#34
default (no effort flag)1955 ±43-85.8$0.0541.0 min#55

Which setting should you use?

For GPT-6 Sol, I'd go straight to max. It sits at #34 on the board, well inside the top half, and at $0.112 a script, the premium over high is small for anything you'll actually publish. The confidence ranges nearly touch, so I'd call the step likely rather than certain. What I'd skip is the default you get with no effort flag: GPT-6 Sol (default) lands 159 Elo below max and saves you only a few cents. If you're churning out drafts in bulk, high is where I'd stop.

Other OpenAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.