Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 (high) leads overall, 88.1 to 83.0.
Grok 4.6 wins this one comfortably. Seventeenth on the board against #45, Elo 2144.5 to 1806.5, and the confidence intervals sit far apart, so the ranking is real. The overall scores tell the same story: 88.14 against 83.05, a 5.09-point gap, and DeepSeek's 9.42 standard deviation means its scripts swing wildly where Grok's 1.43 stays steady. The biggest per-metric gaps are structural: YouTube best practices, 87.08 to 79.86, and numbers handling, 88.24 to 80.85. DeepSeek does take one metric outright, slop scoring, 95.68 to Grok's 95.44, a whisker but a genuine win. Then the ledger. DeepSeek costs about half a cent per script against Grok's about twelve cents, both exact figures, roughly 24x cheaper, and it is open weights. That buys a lot of forgiveness for a 5.09-point gap if you draft at volume and edit anyway.
Pick Grok 4.6 (high) if you want the steadier, higher-scoring script on every run and about twelve cents per draft is a fair price for skipping heavy edits.
Pick DeepSeek V4 Flash 0731 if you draft at volume with an editor in the loop and want open weights at about half a cent per script.
Blue bars: Grok 4.6 (high). Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.
| Grok 4.6 (high) | DeepSeek V4 Flash 0731 | |
|---|---|---|
| Overall / 100 | 88.1 | 83.0 |
| Writing Elo | 2144 | 1806 |
| Run-to-run spread (± overall std) | 1.430 | 9.420 |
| Cost per script (USD) | 0.115 | 0.005 |
| Avg latency (s) | 199.2 | 276.4 |
| Open weights | No | Yes |
Full scorecards: Grok 4.6 (high) · DeepSeek V4 Flash 0731. How scoring works: methodology.
← All comparisons