Towards AITowards AIToneBench

GPT-5.6 Sol (high) vs Gemini 3.1 Pro (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 80.2.

GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
Gemini 3.1 Pro (default)
#63 Elo 1577 · 80.2/100
Cost / script
$0.163 vs $0.118
Human baseline
92.7 GPT-5.6 Sol (high) falls below it · Gemini 3.1 Pro (default) falls below it

The verdict

I use Gemini daily for image generation and deep-research tasks, so no grudge against Google here. But for writing scripts in my voice, Gemini 3.1 Pro at default settings is not the tool, and the numbers are blunt: 80.2 overall against GPT-5.6's 88.38, with Elo intervals that never overlap. Two gaps decide it. Tone, where Gemini sits at 80.0 versus 87.0, so its drafts read reported rather than told. And length adherence, 75.9 versus 91.7, so scripts come back the wrong size and you're trimming or padding by hand. Cost doesn't rescue it either: at twelve cents a script versus fifteen, Gemini is barely cheaper, which removes the usual excuse. Its one real edge is speed, about a minute per script versus over a minute and a half. Not exactly the deciding factor for a writing task. Worth flagging that GPT-5.6 is one of our three judges and rates itself warmly, but it wins with the other two judges as well. So: GPT-5.6 for the scripts, Gemini for the research.

Pick GPT-5.6 Sol if you want drafts that sound spoken and land at the right length for barely more money.
Pick Gemini 3.1 Pro if fast turnaround on rough first drafts matters more to you than voice or length discipline.

Metric by metric

Blue bars: GPT-5.6 Sol (high). Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.3
79.9
Writing Craft & Clarity13% weight
88.1
80.8
Substance, Accuracy & Value15% weight
90.1
80.8
Continuity & Emotion14% weight
86.1
78.4
YouTube Best Practices12% weight
86.5
78.9
Hook Strength10% weight
87.5
82.3
Length Adherence8% weight
91.7
75.9
Slop Score (EQ-Bench + ours)5% weight
93.4
87.7
Visual Cue Quality4% weight
91.0
81.7

Everything else that differs

GPT-5.6 Sol (high)Gemini 3.1 Pro (default)
Overall / 10088.480.2
Writing Elo23151577
Run-to-run spread (± overall std)1.7502.420
Cost per script (USD)0.1630.118
Avg latency (s)106.463.8
Open weightsNoNo

Full scorecards: GPT-5.6 Sol (high) · Gemini 3.1 Pro (default). How scoring works: methodology.

← All comparisons