Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 (high) leads overall, 88.1 to 84.1.
Grok 4.6 (high) wins this pairing without much argument. It sits seventeenth on the board at 2144.5 Elo against #39 and 1847.7 for DeepSeek V4 Pro 0813 (max), a 296.8-point gap, and the confidence intervals do not overlap. The overall scores are closer than the Elo suggests, 88.14 to 84.09, but consistency separates them: DeepSeek's 6.49 standard deviation points to uneven runs where Grok holds steady at 1.43. The per-metric story leans the same way. Cue quality is the widest gap, 86.43 to 79.38, and length adherence follows at 88.26 to 81.21. DeepSeek's case is price and openness. It runs about two cents per script against roughly twelve cents for Grok, close to a 5x saving, and it ships open weights, so it can be self-hosted. That makes it a credible budget option, not a quality peer.
Pick Grok 4.6 (high) if you want the stronger and steadier script on nearly every metric and can absorb roughly 5x the per-script cost.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights and scripts at about two cents each, and can tolerate run-to-run swings in quality.
Blue bars: Grok 4.6 (high). Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.
| Grok 4.6 (high) | DeepSeek V4 Pro 0813 (max) | |
|---|---|---|
| Overall / 100 | 88.1 | 84.1 |
| Writing Elo | 2144 | 1848 |
| Run-to-run spread (± overall std) | 1.430 | 6.490 |
| Cost per script (USD) | 0.115 | 0.022 |
| Avg latency (s) | 199.2 | 227.3 |
| Open weights | No | Yes |
Full scorecards: Grok 4.6 (high) · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.
← All comparisons