Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Qwen3.8 Max leads overall, 83.4 to 82.4. Their confidence intervals overlap, so treat the order as close rather than settled.
Qwen3.8 Max at #38 and GLM-5 at #43, 83.12 against 82.42. Just over a point. Qwen wins on the mechanics that matter for long-form: length adherence 92.90 against 78.61, thirteen points, and anti-slop 87.78 against 75.66, another eleven and a half. GLM answers on the human side, taking voice 83.12 to 81.93, continuity 79.69 to 78.08 and hooks 86.47 to 83.50. So Qwen produces the cleaner, correctly-sized draft and GLM the warmer one. On everything outside quality GLM wins comfortably: two cents a script against 15, open weights against closed, and 117 seconds against 397. Seven times cheaper, three times faster, open. For one point of overall, that is a lot to give up unless the slop and length gaps specifically hurt you.
Pick Qwen3.8 Max if length discipline and slop resistance are the bottleneck and cost is not.
Pick GLM-5 for open weights at a seventh the price and three times the speed, with better voice and hooks.
Blue bars: Qwen3.8 Max. Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.
| Qwen3.8 Max | GLM-5 | |
|---|---|---|
| Overall / 100 | 83.4 | 82.4 |
| Writing Elo | 1866 | 1783 |
| Run-to-run spread (± overall std) | 5.620 | 3.760 |
| Cost per script (USD) | 0.150 | 0.023 |
| Avg latency (s) | 397.4 | 117.3 |
| Open weights | No | Yes |
Full scorecards: Qwen3.8 Max · GLM-5. How scoring works: methodology.
← All comparisons