Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 84.2. Their confidence intervals overlap, so treat the order as close rather than settled.
Grok 4.5 at #30 and DeepSeek V4 Flash 0731 at #34, 85.80 against 84.23. A point and a half, and DeepSeek's wide Elo interval reaches into Grok's, so treat the order as a lean. DeepSeek wins length adherence, 83.57 against 78.67, and their hooks are near identical, 87.96 against 87.55. Grok takes substance 87.42 to 83.48, continuity 83.50 to 82.02, and anti-slop 85.58 to 81.75. So Grok is the more reliable writer and DeepSeek is marginally better at sizing, but neither of these holds a long script especially well. The economics: DeepSeek is half a cent a script against Grok's under four cents, so seven times cheaper, and it is open weights. Grok is much faster, 42 seconds against 289. Grok is the safer pick on writing quality; DeepSeek's case is price, openness, and length control, and it is closing fast.
Pick Grok 4.5 if you want the more reliable draft and near-instant turnaround.
Pick DeepSeek V4 Flash 0731 if open weights or cost matter. Seven times cheaper for two points.
Blue bars: Grok 4.5. Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.
| Grok 4.5 | DeepSeek V4 Flash 0731 | |
|---|---|---|
| Overall / 100 | 85.8 | 84.2 |
| Writing Elo | 2042 | 1955 |
| Run-to-run spread (± overall std) | 2.810 | 9.240 |
| Cost per script (USD) | 0.038 | 0.005 |
| Avg latency (s) | 41.7 | 289.0 |
| Open weights | No | Yes |
Full scorecards: Grok 4.5 · DeepSeek V4 Flash 0731. How scoring works: methodology.
← All comparisons