Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 88.4.
GPT-5.6 Sol at high reasoning is the best non-Anthropic closed model on this board at #17, and it beats Opus 5 on two metrics I care about. Cue quality: 91.02 against 89.56. Anti-slop and number handling: 90.60 against 90.26. If your bottleneck is scripts drifting into AI-flavoured filler or visual cues that don't match the narration, GPT is genuinely the better tool. Everything else goes to Opus, and the overall is 90.80 to 88.38 with 2649.3 Elo against 2315.4. The gap is concentrated in voice and flow. Tone and voice is 91.22 to 87.31, continuity is 90.5 to 86.07. That's a four to five point spread on exactly the thing this benchmark exists to measure, so it matters more than the raw overall suggests. GPT is also more than twice as fast, 106 seconds against 244, and costs 16 cents against 12. Faster, but not cheaper.
Pick Opus 5 (max) if voice fidelity and narrative flow are what you're buying, which for script work is usually the whole point.
Pick GPT-5.6 Sol (high) if you want the cleanest visual cues and the best slop resistance, and you need drafts back in under two minutes.
Blue bars: Claude Opus 5 (max). Orange bars: GPT-5.6 Sol (high). Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | GPT-5.6 Sol (high) | |
|---|---|---|
| Overall / 100 | 90.8 | 88.4 |
| Writing Elo | 2649 | 2315 |
| Run-to-run spread (± overall std) | 1.290 | 1.750 |
| Cost per script (USD) | 0.120 | 0.163 |
| Avg latency (s) | 243.8 | 106.4 |
| Open weights | No | No |
Full scorecards: Claude Opus 5 (max) · GPT-5.6 Sol (high). How scoring works: methodology.
← All comparisons