Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.4 to 85.8.
GPT-5.6 Sol at ultra is #16, Grok 4.5 is #30, and the overall is 88.40 against 85.80. Under three points, but the interesting part is where. Length adherence: 94.34 against 78.67. Seventeen points, the widest gap in this comparison by some distance. Grok does not hit a target word count, and on long-form that alone is a rewrite. Continuity is the other one, 85.72 against 83.50. Grok holds its own on voice, 87.07 against 86.99, and actually edges GPT on hooks, 87.96 to 86.39. So Grok writes decent sentences with good openings that come out the wrong length. The other direction: Grok costs under four cents against GPT's 16, and returns in 42 seconds against 421. That is four times cheaper and eleven times faster. For volume drafting with a human editor, that trade is a lot more defensible than under three points suggests.
Pick GPT-5.6 Sol (ultra) when the draft has to land near a target length without a structural edit.
Pick Grok 4.5 for throughput. Four times cheaper, eleven times faster, decent voice and hooks, and you trim to length yourself.
Blue bars: GPT-5.6 Sol (ultra). Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (ultra) | Grok 4.5 | |
|---|---|---|
| Overall / 100 | 88.4 | 85.8 |
| Writing Elo | 2316 | 2042 |
| Run-to-run spread (± overall std) | 1.970 | 2.810 |
| Cost per script (USD) | 0.137 | 0.038 |
| Avg latency (s) | 420.9 | 41.7 |
| Open weights | No | No |
Full scorecards: GPT-5.6 Sol (ultra) · Grok 4.5. How scoring works: methodology.
← All comparisons