Thinking levels · OpenAI
No, and the setting isn't the problem. gpt-oss 120B (high) and the default version end up 10 Elo apart at #164 and #165, well inside each other's ranges, so turning reasoning up buys nothing you can measure except a longer wait: 1.0 minutes a script instead of 0.4.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Reasoning tokens | Board rank |
|---|---|---|---|---|---|---|---|
| high | 648 ±37 | - | 67.2 | $0.0063 | 1.0 min | 5,537 | #164 |
| default (no effort flag) | 638 ±41 | - | 67.2 | $0.0033 | 0.4 min | 692 | #165 |
The 95% Elo ranges overlap: gpt-oss 120B runs from 598 to 680, and gpt-oss 120B (high) from 606 to 681.
Run it at default. It's faster, and at $0.0033 a script it's the cheaper of the two. The bigger issue is the model itself: its weakest scores are on substance and on the anti-slop check, so you'd be rewriting a lot of what it gives you. It still makes sense if you need open weights you can run yourself, for privacy for example. For scripts, though, an open model like GLM-5.3, at #12, is where I'd start.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.