Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 85.8.
Two closed models, one clear winner, and one real budget question. GPT-5.6 Sol takes it at 88.38 overall against Grok 4.5's 85.8, and the Elo intervals don't overlap, so the ranking holds, even with my standing caveat that GPT-5.6 is one of the three judges and scores itself generously. Here's what surprised me: their hooks land within a fifth of a point of each other, 87.51 for GPT against Grok's 87.96. The separation happens after the opening. Grok scores 78.7 on length adherence against GPT's 91.7, so its scripts miss the target length often enough to need manual cuts, and continuity trails too, 83.5 versus 85.5, sections adjacent rather than connected. Where Grok earns real respect is speed and price: about 42 seconds and 4 cents per script, against 106 seconds and 16 cents. That's roughly an eighth of the cost for a gap of about two and a half points, not thirty. If scripts go out with light editing, GPT-5.6. If you're generating volume and editing everything anyway, Grok's economics are hard to argue with.
Pick GPT-5.6 Sol if scripts ship with light editing and length discipline matters to you.
Pick Grok 4.5 if you're generating volume on a budget and will trim every draft yourself.
Blue bars: GPT-5.6 Sol (high). Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (high) | Grok 4.5 | |
|---|---|---|
| Overall / 100 | 88.4 | 85.8 |
| Writing Elo | 2315 | 2042 |
| Run-to-run spread (± overall std) | 1.750 | 2.810 |
| Cost per script (USD) | 0.163 | 0.038 |
| Avg latency (s) | 106.4 | 41.7 |
| Open weights | No | No |
Full scorecards: GPT-5.6 Sol (high) · Grok 4.5. How scoring works: methodology.
← All comparisons