Towards AITowards AIToneBench

GLM-5.3 vs Grok 4.6

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 86.8 to 86.6. Their confidence intervals overlap, so treat the order as close rather than settled.

GLM-5.3
#24 Elo 2042 · 86.8/100
Grok 4.6
#26 Elo 2035 · 86.6/100
Cost / script
$0.109 vs $0.204
Human baseline
90.2 GLM-5.3 falls below it · Grok 4.6 falls below it

The verdict

GLM-5.3 edges this one on the board, but only just. It sits at #24 with 2041.5 Elo against Grok 4.6's 2035.3 at #26, a 6.2-point gap, and the confidence intervals overlap almost completely, so treat the ranking as a coin flip. Overall it's 86.83 to 86.61, a 0.22-point gap. The interesting part is where each one wins. GLM-5.3 owns the hook: 89.54 against Grok's 80.81, the widest split anywhere on this card. Grok answers on production details, taking visual cues 86.63 to 81.43 and substance 87.53 to 85.93, and it's steadier draft to draft with a 2.53 standard deviation against GLM's 6.6. Tone and voice is a dead heat, 87.89 to 87.86. Then the bill settles it. GLM-5.3 is open weights at about eleven cents per script; Grok 4.6 is closed at about twenty cents. Nearly half the price for the model that's nominally ahead makes this an easy default, unless your scripts live or die on cues and consistency.

Pick GLM-5.3 if you want open weights, the strongest hooks on this card, and a statistically tied result at about eleven cents a script instead of twenty.
Pick Grok 4.6 if you want steadier drafts, better visual cues and substance accuracy, and faster generations, and roughly double the per-script cost doesn't bother you.

Metric by metric

Blue bars: GLM-5.3. Orange bars: Grok 4.6. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.9
87.9
Writing Craft & Clarity13% weight
87.2
86.5
Substance, Accuracy & Value15% weight
85.9
87.5
Continuity & Emotion14% weight
85.1
85.8
YouTube Best Practices12% weight
85.0
86.5
Hook Strength10% weight
89.5
80.8
Length Adherence8% weight
87.5
88.0
Slop Score (EQ-Bench + ours)5% weight
91.6
91.2
Visual Cue Quality4% weight
81.4
86.6

Everything else that differs

GLM-5.3Grok 4.6
Overall / 10086.886.6
Writing Elo20422035
Run-to-run spread (± overall std)6.6002.530
Cost per script (USD)0.1090.204
Avg latency (s)301.9243.1
Open weightsYesNo

Full scorecards: GLM-5.3 · Grok 4.6. How scoring works: methodology.

← All comparisons