Towards AITowards AIToneBench

Claude Opus 5 (max) vs GPT-5.6 Sol (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 88.4.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
Cost / script
$0.120 vs $0.163
Human baseline
92.7 Claude Opus 5 (max) falls below it · GPT-5.6 Sol (high) falls below it

The verdict

GPT-5.6 Sol at high reasoning is the best non-Anthropic closed model on this board at #17, and it beats Opus 5 on two metrics I care about. Cue quality: 91.02 against 89.56. Anti-slop and number handling: 90.60 against 90.26. If your bottleneck is scripts drifting into AI-flavoured filler or visual cues that don't match the narration, GPT is genuinely the better tool. Everything else goes to Opus, and the overall is 90.80 to 88.38 with 2649.3 Elo against 2315.4. The gap is concentrated in voice and flow. Tone and voice is 91.22 to 87.31, continuity is 90.5 to 86.07. That's a four to five point spread on exactly the thing this benchmark exists to measure, so it matters more than the raw overall suggests. GPT is also more than twice as fast, 106 seconds against 244, and costs 16 cents against 12. Faster, but not cheaper.

Pick Opus 5 (max) if voice fidelity and narrative flow are what you're buying, which for script work is usually the whole point.
Pick GPT-5.6 Sol (high) if you want the cleanest visual cues and the best slop resistance, and you need drafts back in under two minutes.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: GPT-5.6 Sol (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
87.3
Writing Craft & Clarity13% weight
91.1
88.1
Substance, Accuracy & Value15% weight
90.6
90.1
Continuity & Emotion14% weight
90.4
86.1
YouTube Best Practices12% weight
89.6
86.5
Hook Strength10% weight
92.0
87.5
Length Adherence8% weight
89.7
91.7
Slop Score (EQ-Bench + ours)5% weight
93.3
93.4
Visual Cue Quality4% weight
89.6
91.0

Everything else that differs

Claude Opus 5 (max)GPT-5.6 Sol (high)
Overall / 10090.888.4
Writing Elo26492315
Run-to-run spread (± overall std)1.2901.750
Cost per script (USD)0.1200.163
Avg latency (s)243.8106.4
Open weightsNoNo

Full scorecards: Claude Opus 5 (max) · GPT-5.6 Sol (high). How scoring works: methodology.

← All comparisons