Towards AITowards AIToneBench

Grok 4.6 vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 leads overall, 86.6 to 82.9.

Grok 4.6
#26 Elo 2035 · 86.6/100
GLM-5
#43 Elo 1748 · 82.9/100
Cost / script
$0.204 vs $0.051
Human baseline
90.2 Grok 4.6 falls below it · GLM-5 falls below it

The verdict

Grok 4.6 wins this one clearly. It sits at #26 with 2035.3 Elo against GLM-5's #43 and 1747.5, and the confidence intervals don't overlap, so the gap is real. On overall score it's 86.61 to 82.88, a 3.73-point gap. Grok takes nearly every metric, and the production ones by the widest margins: length adherence 88.03 vs 80.21, visual cues 86.63 vs 78.89, and numbers-and-slop 87.01 vs 79.39. GLM-5 gets one clean win, hook strength, 86.51 against Grok's 80.81, so its openings land harder even when the rest of the script doesn't. The records tell the same story: 941 wins and 99 losses for Grok against GLM's 692 and 309. Then the budget line. GLM-5 is open weights at about five cents per script against roughly twenty cents for Grok, about 4x cheaper. For a 3.73-point gap, that trade only makes sense if hooks and price matter more to you than everything else.

Pick Grok 4.6 if you want the clearly stronger script across the board, especially on length adherence and visual cues, at about twenty cents per draft.
Pick GLM-5 if you want open weights and the stronger hooks at about 4x less cost per script, and can live with a 3.73-point overall gap.

Metric by metric

Blue bars: Grok 4.6. Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.9
83.4
Writing Craft & Clarity13% weight
86.5
83.7
Substance, Accuracy & Value15% weight
87.5
83.8
Continuity & Emotion14% weight
85.8
80.0
YouTube Best Practices12% weight
86.5
81.7
Hook Strength10% weight
80.8
86.5
Length Adherence8% weight
88.0
80.2
Slop Score (EQ-Bench + ours)5% weight
91.2
87.1
Visual Cue Quality4% weight
86.6
78.9

Everything else that differs

Grok 4.6GLM-5
Overall / 10086.682.9
Writing Elo20351748
Run-to-run spread (± overall std)2.5302.650
Cost per script (USD)0.2040.051
Avg latency (s)243.1188.0
Open weightsNoYes

Full scorecards: Grok 4.6 · GLM-5. How scoring works: methodology.

← All comparisons