Towards AITowards AIToneBench

Grok 4.5 vs DeepSeek V4 Pro (xhigh)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 80.5.

Grok 4.5
#30 Elo 2042 · 85.8/100
DeepSeek V4 Pro (xhigh)
#60 Elo 1601 · 80.5/100
Cost / script
$0.038 vs $0.014
Human baseline
92.7 Grok 4.5 falls below it · DeepSeek V4 Pro (xhigh) falls below it

The verdict

Grok 4.5 takes this one, and at nearly the same price, which makes it the easy call of the seven. Both cost under two cents per script, so cost cancels out and quality decides: Grok wins the overall score with an Elo lead that sits well clear of both confidence intervals. It also comes back in under 30 seconds, against nearly three minutes for DeepSeek, which I notice every time I iterate on drafts. The pattern I see reading them: Grok handles numbers the way I would deliver them on camera, 85.58 versus DeepSeek's 65.96 on that check, where DeepSeek keeps stuffing exact stats into spoken lines. To be fair, DeepSeek beats Grok on one thing that matters to me: hitting the length target, 80.47 against 78.67. Grok tends to run past the word count, and trimming an overlong script is real work. Still, one metric does not flip the fight. Same price, better writing, faster: Grok.

Pick Grok 4.5 if you iterate on drafts and want better writing at the same price with answers in under half a minute.
Pick DeepSeek V4 Pro (xhigh) if hitting the length target out of the box matters more to you than clean stat handling.

Metric by metric

Blue bars: Grok 4.5. Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.1
82.2
Writing Craft & Clarity13% weight
86.6
82.2
Substance, Accuracy & Value15% weight
87.4
81.0
Continuity & Emotion14% weight
83.5
78.5
YouTube Best Practices12% weight
84.4
76.9
Hook Strength10% weight
88.0
84.6
Length Adherence8% weight
78.7
80.5
Slop Score (EQ-Bench + ours)5% weight
90.2
79.4
Visual Cue Quality4% weight
86.5
73.6

Everything else that differs

Grok 4.5DeepSeek V4 Pro (xhigh)
Overall / 10085.880.5
Writing Elo20421601
Run-to-run spread (± overall std)2.8103.670
Cost per script (USD)0.0380.014
Avg latency (s)41.7163.2
Open weightsNoYes

Full scorecards: Grok 4.5 · DeepSeek V4 Pro (xhigh). How scoring works: methodology.

← All comparisons