Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 leads overall, 86.6 to 84.0.
Grok 4.6 wins this one, and the intervals say it's real: Grok sits at #26 with 2035.3 Elo, DeepSeek V4 Pro at #37 with 1880.7, and their confidence intervals don't touch. Overall it's 86.61 against 84.04, a 2.57-point gap. Grok takes almost every metric, and the widest gaps are the production ones: visual cues at 86.63 to 79.47, YouTube best practices at 86.54 to 81.13, and continuity and emotion at 85.78 to 80.77. DeepSeek's one clear win is hook strength, 86.83 against Grok's 80.81, a 6.02-point edge on openings. The records lean the same way, Grok at 941 wins and 99 losses against DeepSeek's 674 and 68 with 628 draws, so judges call a lot of DeepSeek's matchups even. Then the budget line. DeepSeek is open weights at about two cents per script against roughly twenty cents for Grok, nearly 10x cheaper for a 2.57-point gap. If the opening hook matters more than the polish that follows it, that trade is still live.
Pick Grok 4.6 if you want the stronger script almost everywhere, especially on visual cues, continuity, and YouTube structure, and roughly twenty cents per draft is acceptable.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights, the stronger hooks, and about two cents per script, and can live with a 2.57-point overall gap.
Blue bars: Grok 4.6. Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.
| Grok 4.6 | DeepSeek V4 Pro 0813 (max) | |
|---|---|---|
| Overall / 100 | 86.6 | 84.0 |
| Writing Elo | 2035 | 1881 |
| Run-to-run spread (± overall std) | 2.530 | 4.370 |
| Cost per script (USD) | 0.204 | 0.022 |
| Avg latency (s) | 243.1 | 195.2 |
| Open weights | No | Yes |
Full scorecards: Grok 4.6 · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.
← All comparisons