Towards AITowards AIToneBench

GLM-5 vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5 leads overall, 82.9 to 81.1. Their confidence intervals overlap, so treat the order as close rather than settled.

GLM-5
#43 Elo 1748 · 82.9/100
Qwen3.8 Max
#46 Elo 1714 · 81.1/100
Cost / script
$0.051 vs $0.164
Human baseline
90.2 GLM-5 falls below it · Qwen3.8 Max falls below it

The verdict

GLM-5 at #43 and Qwen3.8 Max at #46, 1747.5 Elo against 1713.9, but the confidence intervals overlap, so treat the gap as real but not settled. On overall score it is 82.88 to 81.06, under two points, and GLM is far steadier at 2.65 standard deviation against Qwen's 7.89. GLM wins most of the human side: hooks 86.51 to 83.39, voice 83.39 to 80.65, writing quality 83.7 to 81.05 and continuity 79.97 to 76.22. Qwen answers on mechanics, taking length adherence 86.34 against 80.21 and anti-slop with numbers 86.27 against 79.39, both wide gaps. So GLM produces the stronger, more consistent script and Qwen the correctly-sized, cleaner-with-numbers one. Outside quality GLM also wins comfortably: about five cents per script against about sixteen cents, both exact figures, and open weights against Qwen's closed model. Roughly 3x cheaper, open, and ahead on the board. Qwen only makes sense if the length and numbers gaps specifically hurt you and the cost does not.

Pick Qwen3.8 Max if you need drafts that land on the requested length and keep numbers clean, and the higher per-script cost is not a concern.
Pick GLM-5 if you want open weights, roughly 6x lower cost per script, and stronger hooks and voice, and you can live with looser length control.

Metric by metric

Blue bars: GLM-5. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.4
80.7
Writing Craft & Clarity13% weight
83.7
81.0
Substance, Accuracy & Value15% weight
83.8
81.3
Continuity & Emotion14% weight
80.0
76.2
YouTube Best Practices12% weight
81.7
77.8
Hook Strength10% weight
86.5
83.4
Length Adherence8% weight
80.2
86.3
Slop Score (EQ-Bench + ours)5% weight
87.1
90.4
Visual Cue Quality4% weight
78.9
80.6

Everything else that differs

GLM-5Qwen3.8 Max
Overall / 10082.981.1
Writing Elo17481714
Run-to-run spread (± overall std)2.6507.890
Cost per script (USD)0.0510.164
Avg latency (s)188.0429.0
Open weightsYesNo

Full scorecards: GLM-5 · Qwen3.8 Max. How scoring works: methodology.

← All comparisons