Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4.1 Flash (max) leads overall, 86.8 to 86.6. Their confidence intervals overlap, so treat the order as close rather than settled.
DeepSeek V4.1 Flash (max) has the stronger current result: #22 at 2085.1 Elo and 86.75 overall, compared with Grok 4.6 at #30, 2004.1 Elo, and 86.61 overall. Their 95% Elo confidence intervals overlap, so the exact ordering should be treated as uncertain (2037.20–2122.70 and 1951.50–2050.00). DeepSeek V4.1 Flash (max) has its clearest metric edges in Hook Strength (88.88 versus 80.81) and Length Adherence (89.95 versus 88.03). Grok 4.6 counters on Visual Cue Quality (86.63 versus 82.35) and Substance, Accuracy & Value (87.53 versus 85.23). At the measured run mix, DeepSeek V4.1 Flash (max) costs $0.012 per article versus $0.204 for Grok 4.6; DeepSeek V4.1 Flash (max) is the cheaper route. DeepSeek V4.1 Flash (max) is the open-weights option; Grok 4.6 is closed. On the current automated evidence, DeepSeek V4.1 Flash (max) is the stronger default; Grok 4.6 remains a defensible choice when its specific strengths, price, or deployment profile matter more than the headline rank.
Pick DeepSeek V4.1 Flash (max) when you prioritize the stronger current board result, hook strength, length adherence, open weights, and lower measured cost.
Pick Grok 4.6 when you prioritize visual cue quality, and substance, accuracy & value.
Blue bars: DeepSeek V4.1 Flash (max). Orange bars: Grok 4.6. Same 0–100 scale; the bold bar wins that metric.
| DeepSeek V4.1 Flash (max) | Grok 4.6 | |
|---|---|---|
| Overall / 100 | 86.8 | 86.6 |
| Writing Elo | 2085 | 2004 |
| Run-to-run spread (± overall std) | 6.940 | 2.530 |
| Cost per script (USD) | 0.012 | 0.204 |
| Avg latency (s) | 74.6 | 243.1 |
| Open weights | Yes | No |
Full scorecards: DeepSeek V4.1 Flash (max) · Grok 4.6. How scoring works: methodology.
← All comparisons