Towards AITowards AIToneBench

Claude Opus 5 (max) vs GPT-5.6 Sol (ultra)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 88.4.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
GPT-5.6 Sol (ultra)
#16 Elo 2316 · 88.4/100
Cost / script
$0.120 vs $0.137
Human baseline
92.7 Claude Opus 5 (max) falls below it · GPT-5.6 Sol (ultra) falls below it

The verdict

Ultra is GPT-5.6 Sol's new top rung, and it earns the climb: #16 against xhigh's #18, 88.40 overall against 87.79. Against Opus 5 at max effort it still loses the board, 2649.3 Elo to 2315.9 and 90.80 to 88.40. But look where GPT actually wins, because it is not nothing. Length adherence 94.34 against 90.49, nearly five points, and 94.34 is the best length control of any model I have on the board. Cue quality 91.17 against 89.56. Anti-slop is basically a tie. So if your pain is scripts drifting off the word count or visual cues that do not match the narration, ultra is the better tool, full stop. Opus takes the writing itself: tone 91.22 against 86.99, continuity 90.49 against 85.72, hooks 91.97 against 86.39. That is a four to six point spread on voice, flow and openings, which is the thing this benchmark exists to measure. Cost is close, about 12 cents for Opus against 16 for GPT, and Opus is faster at 244 seconds against 421.

Pick Opus 5 (max) if voice, flow and hooks are what you are buying, which for script work is usually the whole job.
Pick GPT-5.6 Sol (ultra) if hitting an exact length and getting clean visual cues matters more than the last few points of voice.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: GPT-5.6 Sol (ultra). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
87.0
Writing Craft & Clarity13% weight
91.1
88.1
Substance, Accuracy & Value15% weight
90.6
90.5
Continuity & Emotion14% weight
90.4
85.7
YouTube Best Practices12% weight
89.6
86.1
Hook Strength10% weight
92.0
86.4
Length Adherence8% weight
89.7
94.3
Slop Score (EQ-Bench + ours)5% weight
93.3
93.5
Visual Cue Quality4% weight
89.6
91.2

Everything else that differs

Claude Opus 5 (max)GPT-5.6 Sol (ultra)
Overall / 10090.888.4
Writing Elo26492316
Run-to-run spread (± overall std)1.2901.970
Cost per script (USD)0.1200.137
Avg latency (s)243.8420.9
Open weightsNoNo

Full scorecards: Claude Opus 5 (max) · GPT-5.6 Sol (ultra). How scoring works: methodology.

← All comparisons