Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 (high) leads overall, 88.1 to 81.9.
This one is not close. Grok 4.6 (high) sits seventeenth at 2144.5 Elo; GLM-5 is #55 at 1690.6, and their confidence intervals sit hundreds of points apart, so the 453.9-point gap is real. The overalls tell the same story: 88.14 against 81.91, a 6.23 point spread, and GLM-5 is also less consistent, with a standard deviation of 3.59 to Grok's 1.43. The metric detail explains most of it. Grok handles numbers and slop far better, 88.24 vs 75.17 on the anti-slop metric, and it follows length instructions where GLM-5 drifts, 88.26 to 78.66. YouTube craft shows the same pattern, 87.08 against 78.11. What GLM-5 has is the bill and the license. It costs about two cents per script to Grok's about twelve cents, both exact figures, roughly 5x cheaper, and it is open weights, so you can run it yourself. That is a real argument for drafting at volume. For finished scripts, Grok wins comfortably.
Pick Grok 4.6 (high) if you want polished, consistent finished scripts and can absorb a bill of about twelve cents per script.
Pick GLM-5 if you draft at volume, want open weights you can run yourself, and will trade some polish for a script that costs about two cents.
Blue bars: Grok 4.6 (high). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.
| Grok 4.6 (high) | GLM-5 | |
|---|---|---|
| Overall / 100 | 88.1 | 81.9 |
| Writing Elo | 2144 | 1691 |
| Run-to-run spread (± overall std) | 1.430 | 3.590 |
| Cost per script (USD) | 0.115 | 0.024 |
| Avg latency (s) | 199.2 | 121.2 |
| Open weights | No | Yes |
Full scorecards: Grok 4.6 (high) · GLM-5. How scoring works: methodology.
← All comparisons