Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5 leads overall, 82.4 to 81.5.
GLM-5 takes this one, though not as comfortably as the ranks suggest. The Elo intervals don't touch (GLM's floor is 1930, Qwen's ceiling is 1924), so the order is settled even if the margin is thin. The gap that matters is voice and substance: GLM matches my tone at 83.1 vs 80.7 and scores 84.3 vs 81.2 on substance, which in practice is the difference between a script I lightly edit and one I rework. Credit where it's due, Qwen3.7 Max is the more disciplined writer of the two: better length adherence and cleaner number handling, 81.0 against GLM's 77.3, and that 77.3 is GLM's real weakness. But if both need edits either way, I'd rather fix stats than fix voice. The practical side leans the same direction: GLM is open weights and costs roughly a third as much per script, with both figures estimated. Latency is a wash, around two to three minutes each. For a budget or self-hosted pipeline, GLM is the easier pick.
Pick GLM-5 if voice match and open weights at roughly a third of the cost matter more to you than tidy stat handling.
Pick Qwen3.7 Max if you'd rather get disciplined length and cleaner number handling, and open weights aren't a requirement.
Blue bars: GLM-5. Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.
| GLM-5 | Qwen3.7 Max (high) | |
|---|---|---|
| Overall / 100 | 82.4 | 81.5 |
| Writing Elo | 1783 | 1674 |
| Run-to-run spread (± overall std) | 3.760 | 2.580 |
| Cost per script (USD) | 0.023 | 0.054 |
| Avg latency (s) | 117.3 | 157.2 |
| Open weights | Yes | No |
Full scorecards: GLM-5 · Qwen3.7 Max (high). How scoring works: methodology.
← All comparisons