Thinking levels · OpenAI
Barely, and you pay a lot for that barely. The low and max settings end up 18 Elo apart, well inside the noise, and max is the expensive end of that gap: 3.7 times the price, with 8.0 minutes of waiting per script.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Board rank |
|---|---|---|---|---|---|---|
| low | 1926 ±42 | - | 85.6 | $0.049 | 1.5 min | #61 |
| medium | 1908 ±45 | -18 | 85.3 | $0.050 | 1.4 min | #66 |
| high | 1883 ±36 | -25 | 85.4 | $0.061 | 2.0 min | #69 |
| xhigh | 1946 ±44 | +63 | 85.7 | $0.134 | 5.8 min | #56 |
| max | 1944 ±52 | -2 | 85.7 | $0.178 | 8.0 min | #57 |
The 95% Elo ranges overlap: GPT-6.1 Sol (low) runs from 1884 to 1969, and GPT-6.1 Sol (max) from 1887 to 1992.
I'd use low and stop there. The middle settings actually dip a little before xhigh and max climb back, and none of these gaps is big enough to trust. What does move is length adherence: the higher settings land much closer to the target length, and if you trim scripts yourself anyway, that's not worth 3.7 times the price. Honestly, I'd look at GPT-6 Sol (max) first, since the earlier model sits at #34 and scores above every GPT-6.1 Sol setting right now. A bit weird for the newer version, but that's what the judges say.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.