NVIDIA makes the GPUs everyone trains on, but not the writers at the top of this board. 4 configurations of 2 Nemotron models made it, all open weights, and the best of them, Nemotron 3 Ultra 550B, sits at #130. The smaller Super lands much further down, at #162. So this is the cheap, open corner of the board, not where I'd go first for a script.
Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.
| Board rank | Configuration | Elo | Overall | Tone | Craft | Substance | Hook | YouTube | Flow | Slop | Length | Cues | Cost / script | Time / script |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| #130 | Nemotron 3 Ultra 550B · open | 1176 ±56 | 75.8 | 74.9 | 76.3 | 76.9 | 83.4 | 73.5 | 73.2 | 74.8 | 71.3 | 82.0 | $0.017 | 0.4 min |
| #131 | Nemotron 3 Ultra 550B (reasoning) · open | 1175 ±66 | 75.6 | 75.7 | 76.0 | 77.8 | 83.7 | 73.2 | 72.9 | 73.8 | 66.8 | 82.2 | $0.018 | 0.3 min |
| #160 | Nemotron 3 Super 120B (reasoning) · open | 723 ±49 | 68.7 | 67.4 | 70.7 | 72.2 | 75.2 | 64.6 | 67.2 | 70.6 | 62.6 | 65.9 | $0.0027 | 1.3 min |
| #162 | Nemotron 3 Super 120B · open | 665 ±49 | 67.9 | 66.1 | 70.4 | 73.3 | 75.5 | 65.1 | 65.1 | 72.2 | 55.1 | 68.3 | $0.0025 | 1.3 min |
The NVIDIA model I'd use is Nemotron 3 Ultra 550B, at $0.017 a script. There isn't much of a trade-off to weigh. Nemotron 3 Super 120B is cheaper at $0.0025, but you'd be saving pocket change and giving up 510 Elo for it, with the Ultra ahead on nearly every metric we score. I wouldn't bother with the Super for writing.
Zooming out, the Ultra is still a long way down the full board, so I'd only reach for it if I specifically needed open weights from NVIDIA. The main leaderboard has much stronger writers at a similar price.
Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.
Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.
← All comparisons