Towards AITowards AIToneBench

GPT-5.6 Sol (high) vs Grok 4.5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 85.8.

GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
Grok 4.5
#30 Elo 2042 · 85.8/100
Cost / script
$0.163 vs $0.038
Human baseline
92.7 GPT-5.6 Sol (high) falls below it · Grok 4.5 falls below it

The verdict

Two closed models, one clear winner, and one real budget question. GPT-5.6 Sol takes it at 88.38 overall against Grok 4.5's 85.8, and the Elo intervals don't overlap, so the ranking holds, even with my standing caveat that GPT-5.6 is one of the three judges and scores itself generously. Here's what surprised me: their hooks land within a fifth of a point of each other, 87.51 for GPT against Grok's 87.96. The separation happens after the opening. Grok scores 78.7 on length adherence against GPT's 91.7, so its scripts miss the target length often enough to need manual cuts, and continuity trails too, 83.5 versus 85.5, sections adjacent rather than connected. Where Grok earns real respect is speed and price: about 42 seconds and 4 cents per script, against 106 seconds and 16 cents. That's roughly an eighth of the cost for a gap of about two and a half points, not thirty. If scripts go out with light editing, GPT-5.6. If you're generating volume and editing everything anyway, Grok's economics are hard to argue with.

Pick GPT-5.6 Sol if scripts ship with light editing and length discipline matters to you.
Pick Grok 4.5 if you're generating volume on a budget and will trim every draft yourself.

Metric by metric

Blue bars: GPT-5.6 Sol (high). Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.3
87.1
Writing Craft & Clarity13% weight
88.1
86.6
Substance, Accuracy & Value15% weight
90.1
87.4
Continuity & Emotion14% weight
86.1
83.5
YouTube Best Practices12% weight
86.5
84.4
Hook Strength10% weight
87.5
88.0
Length Adherence8% weight
91.7
78.7
Slop Score (EQ-Bench + ours)5% weight
93.4
90.2
Visual Cue Quality4% weight
91.0
86.5

Everything else that differs

GPT-5.6 Sol (high)Grok 4.5
Overall / 10088.485.8
Writing Elo23152042
Run-to-run spread (± overall std)1.7502.810
Cost per script (USD)0.1630.038
Avg latency (s)106.441.7
Open weightsNoNo

Full scorecards: GPT-5.6 Sol (high) · Grok 4.5. How scoring works: methodology.

← All comparisons