Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 80.5.
Grok 4.5 takes this one, and at nearly the same price, which makes it the easy call of the seven. Both cost under two cents per script, so cost cancels out and quality decides: Grok wins the overall score with an Elo lead that sits well clear of both confidence intervals. It also comes back in under 30 seconds, against nearly three minutes for DeepSeek, which I notice every time I iterate on drafts. The pattern I see reading them: Grok handles numbers the way I would deliver them on camera, 85.58 versus DeepSeek's 65.96 on that check, where DeepSeek keeps stuffing exact stats into spoken lines. To be fair, DeepSeek beats Grok on one thing that matters to me: hitting the length target, 80.47 against 78.67. Grok tends to run past the word count, and trimming an overlong script is real work. Still, one metric does not flip the fight. Same price, better writing, faster: Grok.
Pick Grok 4.5 if you iterate on drafts and want better writing at the same price with answers in under half a minute.
Pick DeepSeek V4 Pro (xhigh) if hitting the length target out of the box matters more to you than clean stat handling.
Blue bars: Grok 4.5. Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.
| Grok 4.5 | DeepSeek V4 Pro (xhigh) | |
|---|---|---|
| Overall / 100 | 85.8 | 80.5 |
| Writing Elo | 2042 | 1601 |
| Run-to-run spread (± overall std) | 2.810 | 3.670 |
| Cost per script (USD) | 0.038 | 0.014 |
| Avg latency (s) | 41.7 | 163.2 |
| Open weights | No | Yes |
Full scorecards: Grok 4.5 · DeepSeek V4 Pro (xhigh). How scoring works: methodology.
← All comparisons