Thinking levels · OpenAI
Yes, though almost all of the gain comes from length. GPT-6 Astra (max) scores 84 Elo above the default setting, mostly because it hits the target length far more reliably, while the other scores barely move. It's an expensive fix, too: max costs 3.5 times as much, and a script takes 8.8 minutes where default needs 1.4.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Board rank |
|---|---|---|---|---|---|---|
| max | 1921 ±42 | - | 85.7 | $0.852 | 8.8 min | #62 |
| default (no effort flag) | 1837 ±40 | - | 84.8 | $0.244 | 1.4 min | #72 |
The 95% Elo ranges overlap: GPT-6 Astra (default) runs from 1797 to 1878, and GPT-6 Astra (max) from 1875 to 1959.
To be honest, I wouldn't pick Astra for scripts at either setting. If you're already on it, default is the better deal: max mostly buys you a script that lands closer to length, and you wait several times longer for it. Meanwhile GPT-6 Sol (max) costs $0.112 a script, less than Astra's default, and sits at #34 against #62 for Astra at max. Astra may well earn its price on other work, but on this benchmark, Sol is the GPT-6 I'd write with.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.