Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.5 to 88.1. Their confidence intervals overlap, so treat the order as close rather than settled.
GPT-5.6 Sol (high) takes this one, sitting fourteenth on the board at 2163.5 Elo against Grok 4.6 (high) at seventeenth and 2144.5. The 19-point Elo gap is real but the confidence intervals overlap, so treat the ordering as provisional. Overall scores are nearly identical: 88.47 to 88.14, a 0.33-point spread. The profiles differ more than the totals. GPT-5.6 Sol is the more disciplined drafter, leading on length adherence (91.45 vs 88.26) and cue quality (90.65 vs 86.43), with a substance edge at 89.89 to 87.95. Grok 4.6 counters where the audience meets the script: hook strength (90.03 vs 88.11) and tone (88.75 vs 87.45). Cost tilts toward Grok, at roughly twelve cents per script against about seventeen cents, and its figure is measured rather than estimated. Neither model is open weights. This is a coin flip on quality; format discipline or hooks decides it.
Pick GPT-5.6 Sol (high) if you need scripts that hit the brief, with stronger length adherence, cue quality, and factual substance, and the extra cents per run do not matter.
Pick Grok 4.6 (high) if openings and voice matter most; it hooks harder, reads warmer, and comes in at roughly twelve cents per script with exact rather than estimated cost tracking.
Blue bars: GPT-5.6 Sol (high). Orange bars: Grok 4.6 (high). Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (high) | Grok 4.6 (high) | |
|---|---|---|
| Overall / 100 | 88.5 | 88.1 |
| Writing Elo | 2164 | 2144 |
| Run-to-run spread (± overall std) | 1.720 | 1.430 |
| Cost per script (USD) | 0.168 | 0.115 |
| Avg latency (s) | 107.3 | 199.2 |
| Open weights | No | No |
Full scorecards: GPT-5.6 Sol (high) · Grok 4.6 (high). How scoring works: methodology.
← All comparisons