Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.0 to 86.6.
GPT-5.6 Sol (ultra) wins this one comfortably, and the intervals say it's real: its Elo floor of 2136.6 sits above Grok's ceiling of 2080.8. On the board that's #12 at 2177.4 Elo against #26 at 2035.3, a 142.1-point gap, and 87.97 to 86.61 overall. The clearest split is the hook: 87.69 against 80.81, the widest gap on the card. Length adherence follows at 91.9 to 88.03, and the numbers metric at 90.13 to 87.01. Grok gets one real win, tone and voice at 87.86 to GPT's 86.8, and the two essentially tie on continuity, 85.85 to 85.78. The records point the same way: 1042 wins and 10 losses for GPT-5.6 against Grok's 941 and 99. Then the cost line removes the usual trade. At about fourteen cents a script against about twenty, GPT-5.6 is the cheaper one too. There's no budget case to argue here: the better writer also costs less.
Pick GPT-5.6 Sol (ultra) if you want the stronger script on nearly every metric, hooks especially, at about fourteen cents each — cheaper than the model it beats.
Pick Grok 4.6 if tone and voice is the one metric you weight above everything else — it edges GPT-5.6 there, 87.86 to 86.8 — and you can accept weaker hooks at a higher per-script cost.
Blue bars: GPT-5.6 Sol (ultra). Orange bars: Grok 4.6. Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (ultra) | Grok 4.6 | |
|---|---|---|
| Overall / 100 | 88.0 | 86.6 |
| Writing Elo | 2177 | 2035 |
| Run-to-run spread (± overall std) | 1.450 | 2.530 |
| Cost per script (USD) | 0.143 | 0.204 |
| Avg latency (s) | 195.5 | 243.1 |
| Open weights | No | No |
Full scorecards: GPT-5.6 Sol (ultra) · Grok 4.6. How scoring works: methodology.
← All comparisons