Towards AITowards AIToneBench

GPT-5.6 Sol (ultra) vs Grok 4.5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.4 to 85.8.

GPT-5.6 Sol (ultra)
#16 Elo 2316 · 88.4/100
Grok 4.5
#30 Elo 2042 · 85.8/100
Cost / script
$0.137 vs $0.038
Human baseline
92.7 GPT-5.6 Sol (ultra) falls below it · Grok 4.5 falls below it

The verdict

GPT-5.6 Sol at ultra is #16, Grok 4.5 is #30, and the overall is 88.40 against 85.80. Under three points, but the interesting part is where. Length adherence: 94.34 against 78.67. Seventeen points, the widest gap in this comparison by some distance. Grok does not hit a target word count, and on long-form that alone is a rewrite. Continuity is the other one, 85.72 against 83.50. Grok holds its own on voice, 87.07 against 86.99, and actually edges GPT on hooks, 87.96 to 86.39. So Grok writes decent sentences with good openings that come out the wrong length. The other direction: Grok costs under four cents against GPT's 16, and returns in 42 seconds against 421. That is four times cheaper and eleven times faster. For volume drafting with a human editor, that trade is a lot more defensible than under three points suggests.

Pick GPT-5.6 Sol (ultra) when the draft has to land near a target length without a structural edit.
Pick Grok 4.5 for throughput. Four times cheaper, eleven times faster, decent voice and hooks, and you trim to length yourself.

Metric by metric

Blue bars: GPT-5.6 Sol (ultra). Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.0
87.1
Writing Craft & Clarity13% weight
88.1
86.6
Substance, Accuracy & Value15% weight
90.5
87.4
Continuity & Emotion14% weight
85.7
83.5
YouTube Best Practices12% weight
86.1
84.4
Hook Strength10% weight
86.4
88.0
Length Adherence8% weight
94.3
78.7
Slop Score (EQ-Bench + ours)5% weight
93.5
90.2
Visual Cue Quality4% weight
91.2
86.5

Everything else that differs

GPT-5.6 Sol (ultra)Grok 4.5
Overall / 10088.485.8
Writing Elo23162042
Run-to-run spread (± overall std)1.9702.810
Cost per script (USD)0.1370.038
Avg latency (s)420.941.7
Open weightsNoNo

Full scorecards: GPT-5.6 Sol (ultra) · Grok 4.5. How scoring works: methodology.

← All comparisons