Towards AITowards AIToneBench

Grok 4.6 (high) vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 (high) leads overall, 88.1 to 81.9.

Grok 4.6 (high)
#18 Elo 2144 · 88.1/100
GLM-5
#49 Elo 1691 · 81.9/100
Cost / script
$0.115 vs $0.024
Human baseline
90.3 Grok 4.6 (high) falls below it · GLM-5 falls below it

The verdict

This one is not close. Grok 4.6 (high) sits seventeenth at 2144.5 Elo; GLM-5 is #55 at 1690.6, and their confidence intervals sit hundreds of points apart, so the 453.9-point gap is real. The overalls tell the same story: 88.14 against 81.91, a 6.23 point spread, and GLM-5 is also less consistent, with a standard deviation of 3.59 to Grok's 1.43. The metric detail explains most of it. Grok handles numbers and slop far better, 88.24 vs 75.17 on the anti-slop metric, and it follows length instructions where GLM-5 drifts, 88.26 to 78.66. YouTube craft shows the same pattern, 87.08 against 78.11. What GLM-5 has is the bill and the license. It costs about two cents per script to Grok's about twelve cents, both exact figures, roughly 5x cheaper, and it is open weights, so you can run it yourself. That is a real argument for drafting at volume. For finished scripts, Grok wins comfortably.

Pick Grok 4.6 (high) if you want polished, consistent finished scripts and can absorb a bill of about twelve cents per script.
Pick GLM-5 if you draft at volume, want open weights you can run yourself, and will trade some polish for a script that costs about two cents.

Metric by metric

Blue bars: Grok 4.6 (high). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
88.8
82.5
Writing Craft & Clarity13% weight
88.0
83.5
Substance, Accuracy & Value15% weight
88.0
83.5
Continuity & Emotion14% weight
86.3
78.8
YouTube Best Practices12% weight
87.1
78.1
Hook Strength10% weight
90.0
86.2
Length Adherence8% weight
88.3
78.7
Slop Score (EQ-Bench + ours)5% weight
91.8
85.0
Visual Cue Quality4% weight
86.4
82.0

Everything else that differs

Grok 4.6 (high)GLM-5
Overall / 10088.181.9
Writing Elo21441691
Run-to-run spread (± overall std)1.4303.590
Cost per script (USD)0.1150.024
Avg latency (s)199.2121.2
Open weightsNoYes

Full scorecards: Grok 4.6 (high) · GLM-5. How scoring works: methodology.

← All comparisons