Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 80.2.
Gemini 3.1 Pro sits at #63 here, which surprises people who use it for other work. It is a strong general model that does not do well on this particular job. Against Opus 5 the numbers are blunt: 1577.0 Elo to 2649.3, 80.2 overall to 90.80. More than ten points. The single worst metric is length adherence at 75.91, nearly twelve points below Opus, and continuity at 78.40 is almost as bad. Those two together describe the failure mode exactly: it writes decent paragraphs that don't hold a target length or flow as one piece. Its best metric is cue quality at 81.66, so the visual direction is the part that survives. Cost is basically identical, about 11 cents either way, and Gemini is faster at 60 seconds against 244. But you're paying the same money for a ten point drop, so speed is the only real argument.
Pick Opus 5 (max) unless raw speed is the constraint. At the same price it is not a close call.
Pick Gemini 3.1 Pro if you're already deep in Google tooling and want drafts back in about a minute at the same cost.
Blue bars: Claude Opus 5 (max). Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | Gemini 3.1 Pro (default) | |
|---|---|---|
| Overall / 100 | 90.8 | 80.2 |
| Writing Elo | 2649 | 1577 |
| Run-to-run spread (± overall std) | 1.290 | 2.420 |
| Cost per script (USD) | 0.120 | 0.118 |
| Avg latency (s) | 243.8 | 63.8 |
| Open weights | No | No |
Full scorecards: Claude Opus 5 (max) · Gemini 3.1 Pro (default). How scoring works: methodology.
← All comparisons