Towards AITowards AIToneBench

GPT-5.6 Sol (ultra) vs GLM-5.3

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.0 to 86.8.

GPT-5.6 Sol (ultra)
#12 Elo 2177 · 88.0/100
GLM-5.3
#24 Elo 2042 · 86.8/100
Cost / script
$0.143 vs $0.109
Human baseline
90.2 GPT-5.6 Sol (ultra) falls below it · GLM-5.3 falls below it

The verdict

GPT-5.6 Sol (ultra) wins this one, and the intervals say it plainly: #12 on the board at 2177.4 Elo against GLM-5.3 at #24 with 2041.5, a 135.9-point gap, and the confidence intervals don't touch, so the ranking is real. Overall it's 87.97 to 86.83, a 1.14-point gap, but the consistency number matters more here: GPT-5.6's standard deviation is 1.45 against GLM's 6.6, so GLM swings between strong scripts and weak ones while GPT-5.6 barely moves. Credit where due, GLM-5.3 takes two metrics I care about: tone and voice at 87.89 vs 86.8, and hook strength at 89.54 vs 87.69, so its best openings genuinely land. Everything else goes to GPT-5.6, and some gaps are wide: visual cues at 88.55 vs 81.43, length adherence 91.9 to 87.5, substance 89.01 to 85.93. Now the odd part on cost. GLM is open weights with much cheaper tokens, $1.4 in and $4.4 out per million against $5 and $30, yet per script it lands at about eleven cents against GPT-5.6's fourteen, and it's slower getting there. For roughly three cents more per script, you get the steadier, more complete writer.

Pick GPT-5.6 Sol (ultra) if you want the consistent script every time, with clearly better visual cues, length adherence, and substance, at about fourteen cents each.
Pick GLM-5.3 if you want open weights you can run yourself, stronger hooks and tone on its good days, and can accept scripts that swing more from draft to draft.

Metric by metric

Blue bars: GPT-5.6 Sol (ultra). Orange bars: GLM-5.3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
86.8
87.9
Writing Craft & Clarity13% weight
87.7
87.2
Substance, Accuracy & Value15% weight
89.0
85.9
Continuity & Emotion14% weight
85.8
85.1
YouTube Best Practices12% weight
86.5
85.0
Hook Strength10% weight
87.7
89.5
Length Adherence8% weight
91.9
87.5
Slop Score (EQ-Bench + ours)5% weight
93.3
91.6
Visual Cue Quality4% weight
88.5
81.4

Everything else that differs

GPT-5.6 Sol (ultra)GLM-5.3
Overall / 10088.086.8
Writing Elo21772042
Run-to-run spread (± overall std)1.4506.600
Cost per script (USD)0.1430.109
Avg latency (s)195.5301.9
Open weightsNoYes

Full scorecards: GPT-5.6 Sol (ultra) · GLM-5.3. How scoring works: methodology.

← All comparisons