Towards AITowards AIToneBench

Thinking levels · OpenAI

Does GPT-6.1 Sol write better when it thinks harder?

Barely, and you pay a lot for that barely. The low and max settings end up 18 Elo apart, well inside the noise, and max is the expensive end of that gap: 3.7 times the price, with 8.0 minutes of waiting per script.

X axis

Every setting, step by step

SettingEloStepOverallCost / scriptTime / scriptBoard rank
low1926 ±42-85.6$0.0491.5 min#61
medium1908 ±45-1885.3$0.0501.4 min#66
high1883 ±36-2585.4$0.0612.0 min#69
xhigh1946 ±44+6385.7$0.1345.8 min#56
max1944 ±52-285.7$0.1788.0 min#57

The 95% Elo ranges overlap: GPT-6.1 Sol (low) runs from 1884 to 1969, and GPT-6.1 Sol (max) from 1887 to 1992.

Which setting should you use?

I'd use low and stop there. The middle settings actually dip a little before xhigh and max climb back, and none of these gaps is big enough to trust. What does move is length adherence: the higher settings land much closer to the target length, and if you trim scripts yourself anyway, that's not worth 3.7 times the price. Honestly, I'd look at GPT-6 Sol (max) first, since the earlier model sits at #34 and scores above every GPT-6.1 Sol setting right now. A bit weird for the newer version, but that's what the judges say.

Other OpenAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.