Towards AITowards AIToneBench

GPT-5.6 Sol (ultra) vs Grok 4.6

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.0 to 86.6.

GPT-5.6 Sol (ultra)
#12 Elo 2177 · 88.0/100
Grok 4.6
#26 Elo 2035 · 86.6/100
Cost / script
$0.143 vs $0.204
Human baseline
90.2 GPT-5.6 Sol (ultra) falls below it · Grok 4.6 falls below it

The verdict

GPT-5.6 Sol (ultra) wins this one comfortably, and the intervals say it's real: its Elo floor of 2136.6 sits above Grok's ceiling of 2080.8. On the board that's #12 at 2177.4 Elo against #26 at 2035.3, a 142.1-point gap, and 87.97 to 86.61 overall. The clearest split is the hook: 87.69 against 80.81, the widest gap on the card. Length adherence follows at 91.9 to 88.03, and the numbers metric at 90.13 to 87.01. Grok gets one real win, tone and voice at 87.86 to GPT's 86.8, and the two essentially tie on continuity, 85.85 to 85.78. The records point the same way: 1042 wins and 10 losses for GPT-5.6 against Grok's 941 and 99. Then the cost line removes the usual trade. At about fourteen cents a script against about twenty, GPT-5.6 is the cheaper one too. There's no budget case to argue here: the better writer also costs less.

Pick GPT-5.6 Sol (ultra) if you want the stronger script on nearly every metric, hooks especially, at about fourteen cents each — cheaper than the model it beats.
Pick Grok 4.6 if tone and voice is the one metric you weight above everything else — it edges GPT-5.6 there, 87.86 to 86.8 — and you can accept weaker hooks at a higher per-script cost.

Metric by metric

Blue bars: GPT-5.6 Sol (ultra). Orange bars: Grok 4.6. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
86.8
87.9
Writing Craft & Clarity13% weight
87.7
86.5
Substance, Accuracy & Value15% weight
89.0
87.5
Continuity & Emotion14% weight
85.8
85.8
YouTube Best Practices12% weight
86.5
86.5
Hook Strength10% weight
87.7
80.8
Length Adherence8% weight
91.9
88.0
Slop Score (EQ-Bench + ours)5% weight
93.3
91.2
Visual Cue Quality4% weight
88.5
86.6

Everything else that differs

GPT-5.6 Sol (ultra)Grok 4.6
Overall / 10088.086.6
Writing Elo21772035
Run-to-run spread (± overall std)1.4502.530
Cost per script (USD)0.1430.204
Avg latency (s)195.5243.1
Open weightsNoNo

Full scorecards: GPT-5.6 Sol (ultra) · Grok 4.6. How scoring works: methodology.

← All comparisons