Towards AITowards AIToneBench

GPT-5.6 Sol (high) vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 82.4.

GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
GLM-5
#43 Elo 1783 · 82.4/100
Cost / script
$0.163 vs $0.023
Human baseline
92.7 GPT-5.6 Sol (high) falls below it · GLM-5 falls below it

The verdict

GPT-5.6 takes this one, and the Elo intervals don't even touch, so I'm comfortable calling it. The gap that actually matters is number handling: our anti-slop and numbers check puts GPT-5.6 at 91.0 and GLM-5 at 77.3, the widest split between these two, and GLM's variance there is wild. Some GLM scripts reframe stats the way a human would, others recite them like a quarterly report. Structure is the second miss, a gap of about seven and a half points on YouTube best practices, which in practice means rebuilding retention beats yourself. Now, credit where it's due: GLM-5 is a real open-weights model, its hooks land nearly as hard as GPT's, and at an estimated cent and a half per script versus fifteen cents, it's roughly a tenth of the cost. If you're generating at volume and editing anyway, that math is tempting. One caveat I have to flag: GPT-5.6 sits on our judge panel and scores itself generously. The other two judges still prefer it, just by less. For a publishable draft, GPT-5.6.

Pick GPT-5.6 Sol if you want the closest thing to a publishable draft and can live with fifteen cents per script.
Pick GLM-5 if you want open weights at roughly a tenth of the cost and don't mind rewriting the stats-heavy stretches yourself.

Metric by metric

Blue bars: GPT-5.6 Sol (high). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.3
83.1
Writing Craft & Clarity13% weight
88.1
84.0
Substance, Accuracy & Value15% weight
90.1
84.0
Continuity & Emotion14% weight
86.1
79.7
YouTube Best Practices12% weight
86.5
78.3
Hook Strength10% weight
87.5
86.5
Length Adherence8% weight
91.7
78.6
Slop Score (EQ-Bench + ours)5% weight
93.4
85.3
Visual Cue Quality4% weight
91.0
83.8

Everything else that differs

GPT-5.6 Sol (high)GLM-5
Overall / 10088.482.4
Writing Elo23151783
Run-to-run spread (± overall std)1.7503.760
Cost per script (USD)0.1630.023
Avg latency (s)106.4117.3
Open weightsNoYes

Full scorecards: GPT-5.6 Sol (high) · GLM-5. How scoring works: methodology.

← All comparisons