Towards AITowards AIToneBench

Grok 4.5 vs Gemini 3.1 Pro (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 80.2.

Grok 4.5
#30 Elo 2042 · 85.8/100
Gemini 3.1 Pro (default)
#63 Elo 1577 · 80.2/100
Cost / script
$0.038 vs $0.118
Human baseline
92.7 Grok 4.5 falls below it · Gemini 3.1 Pro (default) falls below it

The verdict

This one is less close than the brand names suggest. Grok 4.5 sits at 85.8 overall against Gemini 3.1 Pro's 80.2, and the Elo gap (2031 vs 1577) is wide enough that the confidence intervals never touch, so this result is not noise. Where does it come from? Mostly voice: Grok matches my tone at 87.1 vs Gemini's 80.0, and that gap is the difference between a script I'd say on camera and one that reads reported. Grok also handles numbers better, 85.6 vs 80.1 on our anti-slop check. They do share one flaw: both hover around 80 on length adherence, so expect drafts that miss the target runtime. Then cost settles whatever was left. Grok runs about 4 cents per script and returns a draft in under 40 seconds; Gemini costs roughly three times more and is slower. Gemini is a strong model in plenty of other places, but as a ghostwriter for my scripts, this is a clear Grok win.

Pick Grok 4.5 if you want the stronger voice match at about 2 cents a script, with drafts back in under 30 seconds.
Pick Gemini 3.1 Pro only if you're locked into the Google stack, since it trails Grok on nearly every metric here while costing about six times more.

Metric by metric

Blue bars: Grok 4.5. Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.1
79.9
Writing Craft & Clarity13% weight
86.6
80.8
Substance, Accuracy & Value15% weight
87.4
80.8
Continuity & Emotion14% weight
83.5
78.4
YouTube Best Practices12% weight
84.4
78.9
Hook Strength10% weight
88.0
82.3
Length Adherence8% weight
78.7
75.9
Slop Score (EQ-Bench + ours)5% weight
90.2
87.7
Visual Cue Quality4% weight
86.5
81.7

Everything else that differs

Grok 4.5Gemini 3.1 Pro (default)
Overall / 10085.880.2
Writing Elo20421577
Run-to-run spread (± overall std)2.8102.420
Cost per script (USD)0.0380.118
Avg latency (s)41.763.8
Open weightsNoNo

Full scorecards: Grok 4.5 · Gemini 3.1 Pro (default). How scoring works: methodology.

← All comparisons