Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5 leads overall, 82.4 to 80.5.
Two open models, both under two cents a script, neither anywhere near the top of my board. That makes it a genuinely fair fight. GLM-5 wins it, though I want to be honest about the margin. The average scores are close, 82.42 against 80.50, with enough spread that they nearly touch. But the rubric-derived Elo tells a cleaner story: GLM's confidence interval sits fully above DeepSeek's across the same scored task set. The gap I can point at is visual cues, 83.79 versus 73.61. GLM gives my editor usable [SHOW:] directions; DeepSeek's read like afterthoughts. Both share the same bad habit of leaving raw stats in the spoken lines, GLM just does it less. At this price neither is a final-draft model for me, they are drafting engines. If I am picking a drafting engine at a cent or two per script, GLM is the one I load first.
Pick GLM-5 if you want the stronger budget open-weights draft, especially for visual cues your editor can actually use.
Pick DeepSeek V4 Pro (xhigh) if you are already running it and a slightly cheaper, roughly comparable draft is enough.
Blue bars: GLM-5. Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.
| GLM-5 | DeepSeek V4 Pro (xhigh) | |
|---|---|---|
| Overall / 100 | 82.4 | 80.5 |
| Writing Elo | 1783 | 1601 |
| Run-to-run spread (± overall std) | 3.760 | 3.670 |
| Cost per script (USD) | 0.023 | 0.014 |
| Avg latency (s) | 117.3 | 163.2 |
| Open weights | Yes | Yes |
Full scorecards: GLM-5 · DeepSeek V4 Pro (xhigh). How scoring works: methodology.
← All comparisons