Towards AITowards AIToneBench

Claude Opus 5 (max) vs Gemini 3.1 Pro (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 80.2.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
Gemini 3.1 Pro (default)
#63 Elo 1577 · 80.2/100
Cost / script
$0.120 vs $0.118
Human baseline
92.7 Claude Opus 5 (max) falls below it · Gemini 3.1 Pro (default) falls below it

The verdict

Gemini 3.1 Pro sits at #63 here, which surprises people who use it for other work. It is a strong general model that does not do well on this particular job. Against Opus 5 the numbers are blunt: 1577.0 Elo to 2649.3, 80.2 overall to 90.80. More than ten points. The single worst metric is length adherence at 75.91, nearly twelve points below Opus, and continuity at 78.40 is almost as bad. Those two together describe the failure mode exactly: it writes decent paragraphs that don't hold a target length or flow as one piece. Its best metric is cue quality at 81.66, so the visual direction is the part that survives. Cost is basically identical, about 11 cents either way, and Gemini is faster at 60 seconds against 244. But you're paying the same money for a ten point drop, so speed is the only real argument.

Pick Opus 5 (max) unless raw speed is the constraint. At the same price it is not a close call.
Pick Gemini 3.1 Pro if you're already deep in Google tooling and want drafts back in about a minute at the same cost.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
79.9
Writing Craft & Clarity13% weight
91.1
80.8
Substance, Accuracy & Value15% weight
90.6
80.8
Continuity & Emotion14% weight
90.4
78.4
YouTube Best Practices12% weight
89.6
78.9
Hook Strength10% weight
92.0
82.3
Length Adherence8% weight
89.7
75.9
Slop Score (EQ-Bench + ours)5% weight
93.3
87.7
Visual Cue Quality4% weight
89.6
81.7

Everything else that differs

Claude Opus 5 (max)Gemini 3.1 Pro (default)
Overall / 10090.880.2
Writing Elo26491577
Run-to-run spread (± overall std)1.2902.420
Cost per script (USD)0.1200.118
Avg latency (s)243.863.8
Open weightsNoNo

Full scorecards: Claude Opus 5 (max) · Gemini 3.1 Pro (default). How scoring works: methodology.

← All comparisons