Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 82.4.
GPT-5.6 takes this one, and the Elo intervals don't even touch, so I'm comfortable calling it. The gap that actually matters is number handling: our anti-slop and numbers check puts GPT-5.6 at 91.0 and GLM-5 at 77.3, the widest split between these two, and GLM's variance there is wild. Some GLM scripts reframe stats the way a human would, others recite them like a quarterly report. Structure is the second miss, a gap of about seven and a half points on YouTube best practices, which in practice means rebuilding retention beats yourself. Now, credit where it's due: GLM-5 is a real open-weights model, its hooks land nearly as hard as GPT's, and at an estimated cent and a half per script versus fifteen cents, it's roughly a tenth of the cost. If you're generating at volume and editing anyway, that math is tempting. One caveat I have to flag: GPT-5.6 sits on our judge panel and scores itself generously. The other two judges still prefer it, just by less. For a publishable draft, GPT-5.6.
Pick GPT-5.6 Sol if you want the closest thing to a publishable draft and can live with fifteen cents per script.
Pick GLM-5 if you want open weights at roughly a tenth of the cost and don't mind rewriting the stats-heavy stretches yourself.
Blue bars: GPT-5.6 Sol (high). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (high) | GLM-5 | |
|---|---|---|
| Overall / 100 | 88.4 | 82.4 |
| Writing Elo | 2315 | 1783 |
| Run-to-run spread (± overall std) | 1.750 | 3.760 |
| Cost per script (USD) | 0.163 | 0.023 |
| Avg latency (s) | 106.4 | 117.3 |
| Open weights | No | Yes |
Full scorecards: GPT-5.6 Sol (high) · GLM-5. How scoring works: methodology.
← All comparisons