Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 88.4.
Ultra is GPT-5.6 Sol's new top rung, and it earns the climb: #16 against xhigh's #18, 88.40 overall against 87.79. Against Opus 5 at max effort it still loses the board, 2649.3 Elo to 2315.9 and 90.80 to 88.40. But look where GPT actually wins, because it is not nothing. Length adherence 94.34 against 90.49, nearly five points, and 94.34 is the best length control of any model I have on the board. Cue quality 91.17 against 89.56. Anti-slop is basically a tie. So if your pain is scripts drifting off the word count or visual cues that do not match the narration, ultra is the better tool, full stop. Opus takes the writing itself: tone 91.22 against 86.99, continuity 90.49 against 85.72, hooks 91.97 against 86.39. That is a four to six point spread on voice, flow and openings, which is the thing this benchmark exists to measure. Cost is close, about 12 cents for Opus against 16 for GPT, and Opus is faster at 244 seconds against 421.
Pick Opus 5 (max) if voice, flow and hooks are what you are buying, which for script work is usually the whole job.
Pick GPT-5.6 Sol (ultra) if hitting an exact length and getting clean visual cues matters more than the last few points of voice.
Blue bars: Claude Opus 5 (max). Orange bars: GPT-5.6 Sol (ultra). Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | GPT-5.6 Sol (ultra) | |
|---|---|---|
| Overall / 100 | 90.8 | 88.4 |
| Writing Elo | 2649 | 2316 |
| Run-to-run spread (± overall std) | 1.290 | 1.970 |
| Cost per script (USD) | 0.120 | 0.137 |
| Avg latency (s) | 243.8 | 420.9 |
| Open weights | No | No |
Full scorecards: Claude Opus 5 (max) · GPT-5.6 Sol (ultra). How scoring works: methodology.
← All comparisons