Towards AITowards AIToneBench

GLM-5 vs Qwen3.7 Max (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5 leads overall, 82.4 to 81.7. Their confidence intervals overlap, so treat the order as close rather than settled.

GLM-5
#43 Elo 1783 · 82.4/100
Qwen3.7 Max (default)
#50 Elo 1699 · 81.7/100
Cost / script
$0.023 vs $0.054
Human baseline
92.7 GLM-5 falls below it · Qwen3.7 Max (default) falls below it

The verdict

GLM-5 at #43 and Qwen3.7 Max at #50, with overalls of 82.42 and 81.73. Under a point apart, and their Elo intervals sit close enough that I'd treat this as a near tie on quality. The differences are in shape, not size. GLM writes with more voice, 83.12 against 81.15, and better hooks, 86.47 against 83.91. Qwen holds length far better, 85.79 against 78.61, and handles slop and numbers slightly better, 77.84 against 75.66. GLM is open weights, Qwen is not, and GLM is less than half the price at just over two cents against five and a half. GLM is also faster, 117 seconds against 157. So on the practical axes that aren't quality, GLM wins most of them: cheaper, faster, open. The one real reason to take Qwen is if you need the length control or you're already on Alibaba tooling.

Pick GLM-5 if you want open weights, half the price, better hooks and more voice.
Pick Qwen3.7 Max if hitting the target length matters more than everything else here.

Metric by metric

Blue bars: GLM-5. Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.1
81.2
Writing Craft & Clarity13% weight
84.0
81.7
Substance, Accuracy & Value15% weight
84.0
80.9
Continuity & Emotion14% weight
79.7
79.0
YouTube Best Practices12% weight
78.3
80.8
Hook Strength10% weight
86.5
84.2
Length Adherence8% weight
78.6
85.8
Slop Score (EQ-Bench + ours)5% weight
85.3
86.0
Visual Cue Quality4% weight
83.8
80.2

Everything else that differs

GLM-5Qwen3.7 Max (default)
Overall / 10082.481.7
Writing Elo17831699
Run-to-run spread (± overall std)3.7603.060
Cost per script (USD)0.0230.054
Avg latency (s)117.3156.8
Open weightsYesNo

Full scorecards: GLM-5 · Qwen3.7 Max (default). How scoring works: methodology.

← All comparisons