Thinking levels · OpenAI
On GPT-5.6 Terra, the effort setting changes your bill a lot more than your script. Going from no reasoning to ultra moves it 16 Elo, and the best score actually comes from xhigh, one step below the top. Ultra, meanwhile, runs $0.159 a script and keeps you waiting 3.8 minutes for each one.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Board rank |
|---|---|---|---|---|---|---|
| no reasoning | 1943 ±42 | - | 85.8 | $0.055 | 1.6 min | #58 |
| low | 1912 ±42 | -31 | 85.4 | $0.059 | 1.0 min | #64 |
| high | 1939 ±44 | +27 | 85.7 | $0.059 | 1.0 min | #59 |
| xhigh | 1976 ±38 | +37 | 86.1 | $0.066 | 1.5 min | #51 |
| ultra | 1959 ±40 | -17 | 85.9 | $0.159 | 3.8 min | #53 |
| default (no effort flag) | 1962 ±39 | - | 86.0 | $0.058 | 1.0 min | #52 |
The 95% Elo ranges overlap: GPT-5.6 Terra (none) runs from 1903 to 1986, and GPT-5.6 Terra (ultra) from 1920 to 2001.
Keep the default. It sits 15 Elo under xhigh, close enough to be noise, and costs $0.058 a script. If you want the best-scoring setting anyway, xhigh adds very little to the price, so go for it. Ultra is the one I'd skip: 2.7 times the default price, and it doesn't score any better than xhigh. The ranges overlap all along this ladder, so on Terra I'd treat the effort knob as a cost knob and leave it alone.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.