Towards AITowards AIToneBench

GPT-5.6 Sol (high) vs Grok 4.6 (high)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.5 to 88.1. Their confidence intervals overlap, so treat the order as close rather than settled.

GPT-5.6 Sol (high)
#15 Elo 2164 · 88.5/100
Grok 4.6 (high)
#18 Elo 2144 · 88.1/100
Cost / script
$0.168 vs $0.115
Human baseline
90.3 GPT-5.6 Sol (high) falls below it · Grok 4.6 (high) falls below it

The verdict

GPT-5.6 Sol (high) takes this one, sitting fourteenth on the board at 2163.5 Elo against Grok 4.6 (high) at seventeenth and 2144.5. The 19-point Elo gap is real but the confidence intervals overlap, so treat the ordering as provisional. Overall scores are nearly identical: 88.47 to 88.14, a 0.33-point spread. The profiles differ more than the totals. GPT-5.6 Sol is the more disciplined drafter, leading on length adherence (91.45 vs 88.26) and cue quality (90.65 vs 86.43), with a substance edge at 89.89 to 87.95. Grok 4.6 counters where the audience meets the script: hook strength (90.03 vs 88.11) and tone (88.75 vs 87.45). Cost tilts toward Grok, at roughly twelve cents per script against about seventeen cents, and its figure is measured rather than estimated. Neither model is open weights. This is a coin flip on quality; format discipline or hooks decides it.

Pick GPT-5.6 Sol (high) if you need scripts that hit the brief, with stronger length adherence, cue quality, and factual substance, and the extra cents per run do not matter.
Pick Grok 4.6 (high) if openings and voice matter most; it hooks harder, reads warmer, and comes in at roughly twelve cents per script with exact rather than estimated cost tracking.

Metric by metric

Blue bars: GPT-5.6 Sol (high). Orange bars: Grok 4.6 (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.5
88.8
Writing Craft & Clarity13% weight
88.3
88.0
Substance, Accuracy & Value15% weight
89.9
88.0
Continuity & Emotion14% weight
86.2
86.3
YouTube Best Practices12% weight
86.7
87.1
Hook Strength10% weight
88.1
90.0
Length Adherence8% weight
91.5
88.3
Slop Score (EQ-Bench + ours)5% weight
93.4
91.8
Visual Cue Quality4% weight
90.7
86.4

Everything else that differs

GPT-5.6 Sol (high)Grok 4.6 (high)
Overall / 10088.588.1
Writing Elo21642144
Run-to-run spread (± overall std)1.7201.430
Cost per script (USD)0.1680.115
Avg latency (s)107.3199.2
Open weightsNoNo

Full scorecards: GPT-5.6 Sol (high) · Grok 4.6 (high). How scoring works: methodology.

← All comparisons