Towards AITowards AIToneBench

Grok 4.6 (high) vs DeepSeek V4 Pro 0813 (max)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 (high) leads overall, 88.1 to 84.1.

Grok 4.6 (high)
#18 Elo 2144 · 88.1/100
DeepSeek V4 Pro 0813 (max)
#38 Elo 1848 · 84.1/100
Cost / script
$0.115 vs $0.022
Human baseline
90.3 Grok 4.6 (high) falls below it · DeepSeek V4 Pro 0813 (max) falls below it

The verdict

Grok 4.6 (high) wins this pairing without much argument. It sits seventeenth on the board at 2144.5 Elo against #39 and 1847.7 for DeepSeek V4 Pro 0813 (max), a 296.8-point gap, and the confidence intervals do not overlap. The overall scores are closer than the Elo suggests, 88.14 to 84.09, but consistency separates them: DeepSeek's 6.49 standard deviation points to uneven runs where Grok holds steady at 1.43. The per-metric story leans the same way. Cue quality is the widest gap, 86.43 to 79.38, and length adherence follows at 88.26 to 81.21. DeepSeek's case is price and openness. It runs about two cents per script against roughly twelve cents for Grok, close to a 5x saving, and it ships open weights, so it can be self-hosted. That makes it a credible budget option, not a quality peer.

Pick Grok 4.6 (high) if you want the stronger and steadier script on nearly every metric and can absorb roughly 5x the per-script cost.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights and scripts at about two cents each, and can tolerate run-to-run swings in quality.

Metric by metric

Blue bars: Grok 4.6 (high). Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
88.8
85.6
Writing Craft & Clarity13% weight
88.0
85.6
Substance, Accuracy & Value15% weight
88.0
83.4
Continuity & Emotion14% weight
86.3
81.8
YouTube Best Practices12% weight
87.1
81.9
Hook Strength10% weight
90.0
88.0
Length Adherence8% weight
88.3
81.2
Slop Score (EQ-Bench + ours)5% weight
91.8
88.8
Visual Cue Quality4% weight
86.4
79.4

Everything else that differs

Grok 4.6 (high)DeepSeek V4 Pro 0813 (max)
Overall / 10088.184.1
Writing Elo21441848
Run-to-run spread (± overall std)1.4306.490
Cost per script (USD)0.1150.022
Avg latency (s)199.2227.3
Open weightsNoYes

Full scorecards: Grok 4.6 (high) · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.

← All comparisons