Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Flash 0731 leads overall, 84.2 to 82.4.
Two open-weights budget models, and the new DeepSeek build takes it cleanly: #34 against #43, 84.23 against 82.42. Just under two points. DeepSeek wins the number discipline decisively, anti-slop 81.75 against 75.66, plus hooks 87.55 to 86.47, voice 85.44 to 83.12, and continuity 82.02 to 79.69. GLM's answers are cue quality, 83.79 against 82.40, and substance, 84.05 against 83.48. On price they are both cheap, but not equally: DeepSeek is half a cent a script against GLM's just over two cents, so four times cheaper still. GLM is faster, 117 seconds against 289. Worth remembering the 0731 build jumped 49 places over the previous DeepSeek Flash, so the trend here is steep.
Pick DeepSeek V4 Flash 0731 as the default budget open model now: cheaper, better voice, better hooks, less slop.
Pick GLM-5 if you want stronger visual cues and substance, or you need the faster turnaround.
Blue bars: DeepSeek V4 Flash 0731. Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.
| DeepSeek V4 Flash 0731 | GLM-5 | |
|---|---|---|
| Overall / 100 | 84.2 | 82.4 |
| Writing Elo | 1955 | 1783 |
| Run-to-run spread (± overall std) | 9.240 | 3.760 |
| Cost per script (USD) | 0.005 | 0.023 |
| Avg latency (s) | 289.0 | 117.3 |
| Open weights | Yes | Yes |
Full scorecards: DeepSeek V4 Flash 0731 · GLM-5. How scoring works: methodology.
← All comparisons