Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5 leads overall, 82.4 to 81.7. Their confidence intervals overlap, so treat the order as close rather than settled.
GLM-5 at #43 and Qwen3.7 Max at #50, with overalls of 82.42 and 81.73. Under a point apart, and their Elo intervals sit close enough that I'd treat this as a near tie on quality. The differences are in shape, not size. GLM writes with more voice, 83.12 against 81.15, and better hooks, 86.47 against 83.91. Qwen holds length far better, 85.79 against 78.61, and handles slop and numbers slightly better, 77.84 against 75.66. GLM is open weights, Qwen is not, and GLM is less than half the price at just over two cents against five and a half. GLM is also faster, 117 seconds against 157. So on the practical axes that aren't quality, GLM wins most of them: cheaper, faster, open. The one real reason to take Qwen is if you need the length control or you're already on Alibaba tooling.
Pick GLM-5 if you want open weights, half the price, better hooks and more voice.
Pick Qwen3.7 Max if hitting the target length matters more than everything else here.
Blue bars: GLM-5. Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.
| GLM-5 | Qwen3.7 Max (default) | |
|---|---|---|
| Overall / 100 | 82.4 | 81.7 |
| Writing Elo | 1783 | 1699 |
| Run-to-run spread (± overall std) | 3.760 | 3.060 |
| Cost per script (USD) | 0.023 | 0.054 |
| Avg latency (s) | 117.3 | 156.8 |
| Open weights | Yes | No |
Full scorecards: GLM-5 · Qwen3.7 Max (default). How scoring works: methodology.
← All comparisons