Towards AITowards AIToneBench

Thinking levels · OpenAI

Does gpt-oss 120B write better when it thinks harder?

No, and the setting isn't the problem. gpt-oss 120B (high) and the default version end up 10 Elo apart at #164 and #165, well inside each other's ranges, so turning reasoning up buys nothing you can measure except a longer wait: 1.0 minutes a script instead of 0.4.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
high648 ±37-67.2$0.00631.0 min5,537#164
default (no effort flag)638 ±41-67.2$0.00330.4 min692#165

The 95% Elo ranges overlap: gpt-oss 120B runs from 598 to 680, and gpt-oss 120B (high) from 606 to 681.

Which setting should you use?

Run it at default. It's faster, and at $0.0033 a script it's the cheaper of the two. The bigger issue is the model itself: its weakest scores are on substance and on the anti-slop check, so you'd be rewriting a lot of what it gives you. It still makes sense if you need open weights you can run yourself, for privacy for example. For scripts, though, an open model like GLM-5.3, at #12, is where I'd start.

Other OpenAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.