Thinking levels · OpenAI
Mostly no, until you hit ultra. GPT-5.6 Sol scores about the same from no reasoning all the way up to xhigh, and only ultra pulls ahead, by 53 Elo. You pay for that jump twice: ultra costs 2.4 times what no reasoning does, and each script takes 3.3 minutes to come back.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Board rank |
|---|---|---|---|---|---|---|
| no reasoning | 2095 ±40 | - | 87.2 | $0.103 | 1.2 min | #37 |
| low | 2091 ±37 | -4 | 87.1 | $0.106 | 1.3 min | #40 |
| high | 2093 ±36 | +2 | 87.2 | $0.111 | 1.4 min | #38 |
| xhigh | 2104 ±34 | +11 | 87.2 | $0.130 | 1.5 min | #35 |
| ultra | 2148 ±35 | +45 | 87.6 | $0.249 | 3.3 min | #31 |
| default (no effort flag) | 2100 ±37 | - | 87.2 | $0.106 | 1.1 min | #36 |
The 95% Elo ranges overlap: GPT-5.6 Sol (none) runs from 2053 to 2134, and GPT-5.6 Sol (ultra) from 2109 to 2180.
One disclosure: GPT-5.6 Sol itself also holds a seat on the judge panel. The panel spans three LLM families plus Jev, so no family scores itself alone, and the methodology explains how that works.
I'd leave it on the default. The Elo barely moves between no reasoning and xhigh, so for scripts, the extra thinking just isn't changing much. Ultra is the only setting that climbs, and it would beat the no reasoning setting in about 58% of script comparisons. That's a lean at best. For one script you really care about, ultra at $0.249 is a cheap bet. At volume, I wouldn't pay 2.3 times the default price for it.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.