Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5 leads overall, 82.9 to 81.1. Their confidence intervals overlap, so treat the order as close rather than settled.
GLM-5 at #43 and Qwen3.8 Max at #46, 1747.5 Elo against 1713.9, but the confidence intervals overlap, so treat the gap as real but not settled. On overall score it is 82.88 to 81.06, under two points, and GLM is far steadier at 2.65 standard deviation against Qwen's 7.89. GLM wins most of the human side: hooks 86.51 to 83.39, voice 83.39 to 80.65, writing quality 83.7 to 81.05 and continuity 79.97 to 76.22. Qwen answers on mechanics, taking length adherence 86.34 against 80.21 and anti-slop with numbers 86.27 against 79.39, both wide gaps. So GLM produces the stronger, more consistent script and Qwen the correctly-sized, cleaner-with-numbers one. Outside quality GLM also wins comfortably: about five cents per script against about sixteen cents, both exact figures, and open weights against Qwen's closed model. Roughly 3x cheaper, open, and ahead on the board. Qwen only makes sense if the length and numbers gaps specifically hurt you and the cost does not.
Pick Qwen3.8 Max if you need drafts that land on the requested length and keep numbers clean, and the higher per-script cost is not a concern.
Pick GLM-5 if you want open weights, roughly 6x lower cost per script, and stronger hooks and voice, and you can live with looser length control.
Blue bars: GLM-5. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| GLM-5 | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 82.9 | 81.1 |
| Writing Elo | 1748 | 1714 |
| Run-to-run spread (± overall std) | 2.650 | 7.890 |
| Cost per script (USD) | 0.051 | 0.164 |
| Avg latency (s) | 188.0 | 429.0 |
| Open weights | Yes | No |
Full scorecards: GLM-5 · Qwen3.8 Max. How scoring works: methodology.
← All comparisons