Towards AITowards AIToneBench

Every NVIDIA model, side by side

NVIDIA makes the GPUs everyone trains on, but not the writers at the top of this board. 4 configurations of 2 Nemotron models made it, all open weights, and the best of them, Nemotron 3 Ultra 550B, sits at #130. The smaller Super lands much further down, at #162. So this is the cheap, open corner of the board, not where I'd go first for a script.

Best writer
Nemotron 3 Ultra 550B#130 · Elo 1176 · $0.017 per script
Cheapest
Nemotron 3 Super 120B#162 · Elo 665 · $0.0025 per script

All 4 configurations

Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.

Board rankConfigurationEloOverallToneCraftSubstanceHookYouTubeFlowSlopLengthCuesCost / scriptTime / script
#130Nemotron 3 Ultra 550B · open1176 ±5675.874.976.376.983.473.573.274.871.382.0$0.0170.4 min
#131Nemotron 3 Ultra 550B (reasoning) · open1175 ±6675.675.776.077.883.773.272.973.866.882.2$0.0180.3 min
#160Nemotron 3 Super 120B (reasoning) · open723 ±4968.767.470.772.275.264.667.270.662.665.9$0.00271.3 min
#162Nemotron 3 Super 120B · open665 ±4967.966.170.473.375.565.165.172.255.168.3$0.00251.3 min

Which one should you use?

The NVIDIA model I'd use is Nemotron 3 Ultra 550B, at $0.017 a script. There isn't much of a trade-off to weigh. Nemotron 3 Super 120B is cheaper at $0.0025, but you'd be saving pocket change and giving up 510 Elo for it, with the Ultra ahead on nearly every metric we score. I wouldn't bother with the Super for writing.

Zooming out, the Ultra is still a long way down the full board, so I'd only reach for it if I specifically needed open weights from NVIDIA. The main leaderboard has much stronger writers at a similar price.

Thinking levels, model by model

Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.

Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.

← All comparisons