Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 86.8 to 81.1.
GLM-5.3 wins this one outright, and there is no angle where Qwen3.8 Max claws it back. GLM sits at #24 with 2041.5 Elo against Qwen's #46 at 1713.9, a 327.6-point gap, and the confidence intervals aren't close to touching, so the ranking is real. Overall it's 86.83 to 81.06, a 5.77-point gap, and GLM sweeps every single metric. The widest split is continuity and emotion, 85.11 vs 76.22, which in practice means Qwen loses the thread of the story more often. Hook strength goes 89.54 to 83.39, tone and voice 87.89 to 80.65. Qwen only stays within about a point on the mechanical stuff: length adherence, 86.34 to GLM's 87.5, and visual cues, 80.59 to 81.43. The records say the same thing, GLM at 866 wins and 8 losses against Qwen's 588 and 258. Then the part that ends the debate: GLM is also cheaper, about eleven cents per script against about sixteen, faster per run, and it's open weights at 400B parameters while Qwen is a closed trillion-parameter model. Better, cheaper, and open. There is no trade to weigh here.
Pick GLM-5.3 if you want the stronger script on every metric, open weights at 400B parameters you can run yourself, and the lower bill at about eleven cents per draft.
Pick Qwen3.8 Max only if you're already committed to its hosted API and the two metrics where it nearly keeps pace, length adherence and visual cues, are the ones you care about most.
Blue bars: GLM-5.3. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| GLM-5.3 | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 86.8 | 81.1 |
| Writing Elo | 2042 | 1714 |
| Run-to-run spread (± overall std) | 6.600 | 7.890 |
| Cost per script (USD) | 0.109 | 0.164 |
| Avg latency (s) | 301.9 | 429.0 |
| Open weights | Yes | No |
Full scorecards: GLM-5.3 · Qwen3.8 Max. How scoring works: methodology.
← All comparisons