Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 80.2.
I use Gemini daily for image generation and deep-research tasks, so no grudge against Google here. But for writing scripts in my voice, Gemini 3.1 Pro at default settings is not the tool, and the numbers are blunt: 80.2 overall against GPT-5.6's 88.38, with Elo intervals that never overlap. Two gaps decide it. Tone, where Gemini sits at 80.0 versus 87.0, so its drafts read reported rather than told. And length adherence, 75.9 versus 91.7, so scripts come back the wrong size and you're trimming or padding by hand. Cost doesn't rescue it either: at twelve cents a script versus fifteen, Gemini is barely cheaper, which removes the usual excuse. Its one real edge is speed, about a minute per script versus over a minute and a half. Not exactly the deciding factor for a writing task. Worth flagging that GPT-5.6 is one of our three judges and rates itself warmly, but it wins with the other two judges as well. So: GPT-5.6 for the scripts, Gemini for the research.
Pick GPT-5.6 Sol if you want drafts that sound spoken and land at the right length for barely more money.
Pick Gemini 3.1 Pro if fast turnaround on rough first drafts matters more to you than voice or length discipline.
Blue bars: GPT-5.6 Sol (high). Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (high) | Gemini 3.1 Pro (default) | |
|---|---|---|
| Overall / 100 | 88.4 | 80.2 |
| Writing Elo | 2315 | 1577 |
| Run-to-run spread (± overall std) | 1.750 | 2.420 |
| Cost per script (USD) | 0.163 | 0.118 |
| Avg latency (s) | 106.4 | 63.8 |
| Open weights | No | No |
Full scorecards: GPT-5.6 Sol (high) · Gemini 3.1 Pro (default). How scoring works: methodology.
← All comparisons