Towards AITowards AIToneBench

Qwen3.8 Max vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Qwen3.8 Max leads overall, 83.4 to 82.4. Their confidence intervals overlap, so treat the order as close rather than settled.

Qwen3.8 Max
#38 Elo 1866 · 83.4/100
GLM-5
#43 Elo 1783 · 82.4/100
Cost / script
$0.150 vs $0.023
Human baseline
92.7 Qwen3.8 Max falls below it · GLM-5 falls below it

The verdict

Qwen3.8 Max at #38 and GLM-5 at #43, 83.12 against 82.42. Just over a point. Qwen wins on the mechanics that matter for long-form: length adherence 92.90 against 78.61, thirteen points, and anti-slop 87.78 against 75.66, another eleven and a half. GLM answers on the human side, taking voice 83.12 to 81.93, continuity 79.69 to 78.08 and hooks 86.47 to 83.50. So Qwen produces the cleaner, correctly-sized draft and GLM the warmer one. On everything outside quality GLM wins comfortably: two cents a script against 15, open weights against closed, and 117 seconds against 397. Seven times cheaper, three times faster, open. For one point of overall, that is a lot to give up unless the slop and length gaps specifically hurt you.

Pick Qwen3.8 Max if length discipline and slop resistance are the bottleneck and cost is not.
Pick GLM-5 for open weights at a seventh the price and three times the speed, with better voice and hooks.

Metric by metric

Blue bars: Qwen3.8 Max. Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
81.9
83.1
Writing Craft & Clarity13% weight
83.4
84.0
Substance, Accuracy & Value15% weight
84.4
84.0
Continuity & Emotion14% weight
78.1
79.7
YouTube Best Practices12% weight
80.4
78.3
Hook Strength10% weight
83.5
86.5
Length Adherence8% weight
92.9
78.6
Slop Score (EQ-Bench + ours)5% weight
92.0
85.3
Visual Cue Quality4% weight
85.2
83.8

Everything else that differs

Qwen3.8 MaxGLM-5
Overall / 10083.482.4
Writing Elo18661783
Run-to-run spread (± overall std)5.6203.760
Cost per script (USD)0.1500.023
Avg latency (s)397.4117.3
Open weightsNoYes

Full scorecards: Qwen3.8 Max · GLM-5. How scoring works: methodology.

← All comparisons