Towards AITowards AIToneBench

GLM-5 vs Qwen3.7 Max (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5 leads overall, 82.4 to 81.5.

GLM-5
#43 Elo 1783 · 82.4/100
Qwen3.7 Max (high)
#55 Elo 1674 · 81.5/100
Cost / script
$0.023 vs $0.054
Human baseline
92.7 GLM-5 falls below it · Qwen3.7 Max (high) falls below it

The verdict

GLM-5 takes this one, though not as comfortably as the ranks suggest. The Elo intervals don't touch (GLM's floor is 1930, Qwen's ceiling is 1924), so the order is settled even if the margin is thin. The gap that matters is voice and substance: GLM matches my tone at 83.1 vs 80.7 and scores 84.3 vs 81.2 on substance, which in practice is the difference between a script I lightly edit and one I rework. Credit where it's due, Qwen3.7 Max is the more disciplined writer of the two: better length adherence and cleaner number handling, 81.0 against GLM's 77.3, and that 77.3 is GLM's real weakness. But if both need edits either way, I'd rather fix stats than fix voice. The practical side leans the same direction: GLM is open weights and costs roughly a third as much per script, with both figures estimated. Latency is a wash, around two to three minutes each. For a budget or self-hosted pipeline, GLM is the easier pick.

Pick GLM-5 if voice match and open weights at roughly a third of the cost matter more to you than tidy stat handling.
Pick Qwen3.7 Max if you'd rather get disciplined length and cleaner number handling, and open weights aren't a requirement.

Metric by metric

Blue bars: GLM-5. Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.1
80.9
Writing Craft & Clarity13% weight
84.0
81.5
Substance, Accuracy & Value15% weight
84.0
81.2
Continuity & Emotion14% weight
79.7
79.2
YouTube Best Practices12% weight
78.3
80.7
Hook Strength10% weight
86.5
85.0
Length Adherence8% weight
78.6
83.3
Slop Score (EQ-Bench + ours)5% weight
85.3
85.9
Visual Cue Quality4% weight
83.8
78.4

Everything else that differs

GLM-5Qwen3.7 Max (high)
Overall / 10082.481.5
Writing Elo17831674
Run-to-run spread (± overall std)3.7602.580
Cost per script (USD)0.0230.054
Avg latency (s)117.3157.2
Open weightsYesNo

Full scorecards: GLM-5 · Qwen3.7 Max (high). How scoring works: methodology.

← All comparisons