Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 86.8 to 86.6. Their confidence intervals overlap, so treat the order as close rather than settled.
GLM-5.3 edges this one on the board, but only just. It sits at #24 with 2041.5 Elo against Grok 4.6's 2035.3 at #26, a 6.2-point gap, and the confidence intervals overlap almost completely, so treat the ranking as a coin flip. Overall it's 86.83 to 86.61, a 0.22-point gap. The interesting part is where each one wins. GLM-5.3 owns the hook: 89.54 against Grok's 80.81, the widest split anywhere on this card. Grok answers on production details, taking visual cues 86.63 to 81.43 and substance 87.53 to 85.93, and it's steadier draft to draft with a 2.53 standard deviation against GLM's 6.6. Tone and voice is a dead heat, 87.89 to 87.86. Then the bill settles it. GLM-5.3 is open weights at about eleven cents per script; Grok 4.6 is closed at about twenty cents. Nearly half the price for the model that's nominally ahead makes this an easy default, unless your scripts live or die on cues and consistency.
Pick GLM-5.3 if you want open weights, the strongest hooks on this card, and a statistically tied result at about eleven cents a script instead of twenty.
Pick Grok 4.6 if you want steadier drafts, better visual cues and substance accuracy, and faster generations, and roughly double the per-script cost doesn't bother you.
Blue bars: GLM-5.3. Orange bars: Grok 4.6. Same 0–100 scale; the bold bar wins that metric.
| GLM-5.3 | Grok 4.6 | |
|---|---|---|
| Overall / 100 | 86.8 | 86.6 |
| Writing Elo | 2042 | 2035 |
| Run-to-run spread (± overall std) | 6.600 | 2.530 |
| Cost per script (USD) | 0.109 | 0.204 |
| Avg latency (s) | 301.9 | 243.1 |
| Open weights | Yes | No |
Full scorecards: GLM-5.3 · Grok 4.6. How scoring works: methodology.
← All comparisons