Towards AITowards AIToneBench

Grok 4.6 vs DeepSeek V4 Pro 0813 (max)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 leads overall, 86.6 to 84.0.

Grok 4.6
#26 Elo 2035 · 86.6/100
DeepSeek V4 Pro 0813 (max)
#37 Elo 1881 · 84.0/100
Cost / script
$0.204 vs $0.022
Human baseline
90.2 Grok 4.6 falls below it · DeepSeek V4 Pro 0813 (max) falls below it

The verdict

Grok 4.6 wins this one, and the intervals say it's real: Grok sits at #26 with 2035.3 Elo, DeepSeek V4 Pro at #37 with 1880.7, and their confidence intervals don't touch. Overall it's 86.61 against 84.04, a 2.57-point gap. Grok takes almost every metric, and the widest gaps are the production ones: visual cues at 86.63 to 79.47, YouTube best practices at 86.54 to 81.13, and continuity and emotion at 85.78 to 80.77. DeepSeek's one clear win is hook strength, 86.83 against Grok's 80.81, a 6.02-point edge on openings. The records lean the same way, Grok at 941 wins and 99 losses against DeepSeek's 674 and 68 with 628 draws, so judges call a lot of DeepSeek's matchups even. Then the budget line. DeepSeek is open weights at about two cents per script against roughly twenty cents for Grok, nearly 10x cheaper for a 2.57-point gap. If the opening hook matters more than the polish that follows it, that trade is still live.

Pick Grok 4.6 if you want the stronger script almost everywhere, especially on visual cues, continuity, and YouTube structure, and roughly twenty cents per draft is acceptable.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights, the stronger hooks, and about two cents per script, and can live with a 2.57-point overall gap.

Metric by metric

Blue bars: Grok 4.6. Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.9
84.7
Writing Craft & Clarity13% weight
86.5
84.9
Substance, Accuracy & Value15% weight
87.5
84.7
Continuity & Emotion14% weight
85.8
80.8
YouTube Best Practices12% weight
86.5
81.1
Hook Strength10% weight
80.8
86.8
Length Adherence8% weight
88.0
86.2
Slop Score (EQ-Bench + ours)5% weight
91.2
88.0
Visual Cue Quality4% weight
86.6
79.5

Everything else that differs

Grok 4.6DeepSeek V4 Pro 0813 (max)
Overall / 10086.684.0
Writing Elo20351881
Run-to-run spread (± overall std)2.5304.370
Cost per script (USD)0.2040.022
Avg latency (s)243.1195.2
Open weightsNoYes

Full scorecards: Grok 4.6 · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.

← All comparisons