Towards AITowards AIToneBench

DeepSeek V4 Flash 0731 vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Flash 0731 leads overall, 84.2 to 82.4.

DeepSeek V4 Flash 0731
#34 Elo 1955 · 84.2/100
GLM-5
#43 Elo 1783 · 82.4/100
Cost / script
$0.005 vs $0.023
Human baseline
92.7 DeepSeek V4 Flash 0731 falls below it · GLM-5 falls below it

The verdict

Two open-weights budget models, and the new DeepSeek build takes it cleanly: #34 against #43, 84.23 against 82.42. Just under two points. DeepSeek wins the number discipline decisively, anti-slop 81.75 against 75.66, plus hooks 87.55 to 86.47, voice 85.44 to 83.12, and continuity 82.02 to 79.69. GLM's answers are cue quality, 83.79 against 82.40, and substance, 84.05 against 83.48. On price they are both cheap, but not equally: DeepSeek is half a cent a script against GLM's just over two cents, so four times cheaper still. GLM is faster, 117 seconds against 289. Worth remembering the 0731 build jumped 49 places over the previous DeepSeek Flash, so the trend here is steep.

Pick DeepSeek V4 Flash 0731 as the default budget open model now: cheaper, better voice, better hooks, less slop.
Pick GLM-5 if you want stronger visual cues and substance, or you need the faster turnaround.

Metric by metric

Blue bars: DeepSeek V4 Flash 0731. Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
85.4
83.1
Writing Craft & Clarity13% weight
85.3
84.0
Substance, Accuracy & Value15% weight
83.5
84.0
Continuity & Emotion14% weight
82.0
79.7
YouTube Best Practices12% weight
81.1
78.3
Hook Strength10% weight
87.5
86.5
Length Adherence8% weight
83.6
78.6
Slop Score (EQ-Bench + ours)5% weight
88.8
85.3
Visual Cue Quality4% weight
82.4
83.8

Everything else that differs

DeepSeek V4 Flash 0731GLM-5
Overall / 10084.282.4
Writing Elo19551783
Run-to-run spread (± overall std)9.2403.760
Cost per script (USD)0.0050.023
Avg latency (s)289.0117.3
Open weightsYesYes

Full scorecards: DeepSeek V4 Flash 0731 · GLM-5. How scoring works: methodology.

← All comparisons