Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 80.2.
This one is less close than the brand names suggest. Grok 4.5 sits at 85.8 overall against Gemini 3.1 Pro's 80.2, and the Elo gap (2031 vs 1577) is wide enough that the confidence intervals never touch, so this result is not noise. Where does it come from? Mostly voice: Grok matches my tone at 87.1 vs Gemini's 80.0, and that gap is the difference between a script I'd say on camera and one that reads reported. Grok also handles numbers better, 85.6 vs 80.1 on our anti-slop check. They do share one flaw: both hover around 80 on length adherence, so expect drafts that miss the target runtime. Then cost settles whatever was left. Grok runs about 4 cents per script and returns a draft in under 40 seconds; Gemini costs roughly three times more and is slower. Gemini is a strong model in plenty of other places, but as a ghostwriter for my scripts, this is a clear Grok win.
Pick Grok 4.5 if you want the stronger voice match at about 2 cents a script, with drafts back in under 30 seconds.
Pick Gemini 3.1 Pro only if you're locked into the Google stack, since it trails Grok on nearly every metric here while costing about six times more.
Blue bars: Grok 4.5. Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.
| Grok 4.5 | Gemini 3.1 Pro (default) | |
|---|---|---|
| Overall / 100 | 85.8 | 80.2 |
| Writing Elo | 2042 | 1577 |
| Run-to-run spread (± overall std) | 2.810 | 2.420 |
| Cost per script (USD) | 0.038 | 0.118 |
| Avg latency (s) | 41.7 | 63.8 |
| Open weights | No | No |
Full scorecards: Grok 4.5 · Gemini 3.1 Pro (default). How scoring works: methodology.
← All comparisons