Towards AITowards AIToneBench

Thinking levels · NVIDIA

Does Nemotron 3 Super 120B write better when it thinks harder?

A little, maybe. Setting reasoning to high moves Nemotron 3 Super 120B up 57 Elo over its default, but the two ranges still overlap, so I wouldn't call it a proven gain. What's at stake is tiny anyway: both settings cost a fraction of a cent per script.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
high723 ±49-68.7$0.00271.3 min2,349#160
default (no effort flag)665 ±49-67.9$0.00251.3 min1,924#162

The 95% Elo ranges overlap: Nemotron 3 Super 120B runs from 623 to 721, and Nemotron 3 Super 120B (reasoning) from 679 to 777.

Which setting should you use?

If I had to run it, I'd use Nemotron 3 Super 120B (reasoning). It's ahead of the default on Elo, it costs $0.0027 against $0.0025, and both take about 1.3 minutes, so leaving reasoning on has no real downside. Treat that as a practical lean, not a measured win, since a rerun could erase a step like this. Reasoning on or off, you're still down around #160, so I wouldn't start here for scripts. If you want to stay with NVIDIA, the bigger Nemotron 3 Ultra 550B scores well above both at #130 for $0.017 per script.

Other NVIDIA ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every NVIDIA model side by side.