Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 leads overall, 86.6 to 82.9.
Grok 4.6 wins this one clearly. It sits at #26 with 2035.3 Elo against GLM-5's #43 and 1747.5, and the confidence intervals don't overlap, so the gap is real. On overall score it's 86.61 to 82.88, a 3.73-point gap. Grok takes nearly every metric, and the production ones by the widest margins: length adherence 88.03 vs 80.21, visual cues 86.63 vs 78.89, and numbers-and-slop 87.01 vs 79.39. GLM-5 gets one clean win, hook strength, 86.51 against Grok's 80.81, so its openings land harder even when the rest of the script doesn't. The records tell the same story: 941 wins and 99 losses for Grok against GLM's 692 and 309. Then the budget line. GLM-5 is open weights at about five cents per script against roughly twenty cents for Grok, about 4x cheaper. For a 3.73-point gap, that trade only makes sense if hooks and price matter more to you than everything else.
Pick Grok 4.6 if you want the clearly stronger script across the board, especially on length adherence and visual cues, at about twenty cents per draft.
Pick GLM-5 if you want open weights and the stronger hooks at about 4x less cost per script, and can live with a 3.73-point overall gap.
Blue bars: Grok 4.6. Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.
| Grok 4.6 | GLM-5 | |
|---|---|---|
| Overall / 100 | 86.6 | 82.9 |
| Writing Elo | 2035 | 1748 |
| Run-to-run spread (± overall std) | 2.530 | 2.650 |
| Cost per script (USD) | 0.204 | 0.051 |
| Avg latency (s) | 243.1 | 188.0 |
| Open weights | No | Yes |
Full scorecards: Grok 4.6 · GLM-5. How scoring works: methodology.
← All comparisons