Towards AITowards AIToneBench

Grok 4.5 vs DeepSeek V4 Flash 0731

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 84.2. Their confidence intervals overlap, so treat the order as close rather than settled.

Grok 4.5
#30 Elo 2042 · 85.8/100
DeepSeek V4 Flash 0731
#34 Elo 1955 · 84.2/100
Cost / script
$0.038 vs $0.005
Human baseline
92.7 Grok 4.5 falls below it · DeepSeek V4 Flash 0731 falls below it

The verdict

Grok 4.5 at #30 and DeepSeek V4 Flash 0731 at #34, 85.80 against 84.23. A point and a half, and DeepSeek's wide Elo interval reaches into Grok's, so treat the order as a lean. DeepSeek wins length adherence, 83.57 against 78.67, and their hooks are near identical, 87.96 against 87.55. Grok takes substance 87.42 to 83.48, continuity 83.50 to 82.02, and anti-slop 85.58 to 81.75. So Grok is the more reliable writer and DeepSeek is marginally better at sizing, but neither of these holds a long script especially well. The economics: DeepSeek is half a cent a script against Grok's under four cents, so seven times cheaper, and it is open weights. Grok is much faster, 42 seconds against 289. Grok is the safer pick on writing quality; DeepSeek's case is price, openness, and length control, and it is closing fast.

Pick Grok 4.5 if you want the more reliable draft and near-instant turnaround.
Pick DeepSeek V4 Flash 0731 if open weights or cost matter. Seven times cheaper for two points.

Metric by metric

Blue bars: Grok 4.5. Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.1
85.4
Writing Craft & Clarity13% weight
86.6
85.3
Substance, Accuracy & Value15% weight
87.4
83.5
Continuity & Emotion14% weight
83.5
82.0
YouTube Best Practices12% weight
84.4
81.1
Hook Strength10% weight
88.0
87.5
Length Adherence8% weight
78.7
83.6
Slop Score (EQ-Bench + ours)5% weight
90.2
88.8
Visual Cue Quality4% weight
86.5
82.4

Everything else that differs

Grok 4.5DeepSeek V4 Flash 0731
Overall / 10085.884.2
Writing Elo20421955
Run-to-run spread (± overall std)2.8109.240
Cost per script (USD)0.0380.005
Avg latency (s)41.7289.0
Open weightsNoYes

Full scorecards: Grok 4.5 · DeepSeek V4 Flash 0731. How scoring works: methodology.

← All comparisons