Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 86.8 to 84.0.
GLM-5.3 wins this one, and the intervals say it's real: GLM's Elo floor of 1989.0 sits above DeepSeek's ceiling of 1946.4. GLM ranks #24 at 2041.5 Elo against DeepSeek V4 Pro's #37 at 1880.7, a 160.8-point gap, and takes the overall score 86.83 to 84.04. It's a sweep across every metric, and the biggest gaps land where a script stops sounding like a script: numbers-and-slop at 86.75 vs 82.09, continuity and emotion at 85.11 vs 80.77, and YouTube best practices at 85.02 vs 81.13. The records agree, 866 wins and 8 losses for GLM against DeepSeek's 674 and 68. Both are open weights, so the tiebreaker is the bill. DeepSeek runs about two cents per script against GLM's roughly eleven, about 5x cheaper, and it's faster too. A 2.79-point overall gap at 5x less cost is a live trade if the draft gets rewritten anyway. If it ships close to as-written, GLM earns the premium.
Pick GLM-5.3 if you want the clear board winner, with stronger continuity, hooks, and fewer slop tells, at roughly eleven cents a script.
Pick DeepSeek V4 Pro 0813 (max) if you want drafts at about two cents each, delivered faster, and can accept a 2.79-point overall gap on scripts you plan to edit anyway.
Blue bars: GLM-5.3. Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.
| GLM-5.3 | DeepSeek V4 Pro 0813 (max) | |
|---|---|---|
| Overall / 100 | 86.8 | 84.0 |
| Writing Elo | 2042 | 1881 |
| Run-to-run spread (± overall std) | 6.600 | 4.370 |
| Cost per script (USD) | 0.109 | 0.022 |
| Avg latency (s) | 301.9 | 195.2 |
| Open weights | Yes | Yes |
Full scorecards: GLM-5.3 · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.
← All comparisons