Towards AITowards AIToneBench

GLM-5.3 vs GPT-5.6 Sol (ultra)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 88.3 to 88.0. Their confidence intervals overlap, so treat the order as close rather than settled.

GLM-5.3
#7 Elo 2256 · 88.3/100
GPT-5.6 Sol (ultra)
#19 Elo 2218 · 88.0/100
Cost / script
$0.114 vs $0.363
Human baseline
90.2 GLM-5.3 falls below it · GPT-5.6 Sol (ultra) falls below it

The verdict

This one is closer than the ranks suggest. GLM-5.3 sits at #7 with 2256.0 Elo and 88.34 overall; GPT-5.6 Sol (ultra) is at #19 with 2218.2 Elo and 87.97 overall. Their 95% Elo intervals overlap (2221.7 to 2291.5 against 2177.5 to 2256.6), so do not read much into the exact order. Where GLM-5.3 pulls ahead: Hook Strength (89.95 vs 87.69) and Tone & Voice Match (88.72 vs 86.80). GPT-5.6 Sol (ultra) still wins on Visual Cue Quality (88.55 vs 83.02) and Length Adherence (91.90 vs 90.43), so it is not a clean sweep. Price points the same way: GLM-5.3 costs about $0.114 per article against $0.363 for GPT-5.6 Sol (ultra). GLM-5.3 publishes open weights; GPT-5.6 Sol (ultra) does not. Either is a reasonable default: lean GLM-5.3 for hook strength, voice match, the lower price, and open weights, GPT-5.6 Sol (ultra) for visual cues and length adherence.

Pick GPT-5.6 Sol (ultra) for visual cues and length adherence.
Pick GLM-5.3 for the stronger board result, hook strength, voice match, open weights, and the lower price.

Metric by metric

Blue bars: GLM-5.3. Orange bars: GPT-5.6 Sol (ultra). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
88.7
86.8
Writing Craft & Clarity13% weight
88.2
87.7
Substance, Accuracy & Value15% weight
87.9
89.0
Continuity & Emotion14% weight
87.0
85.8
YouTube Best Practices12% weight
87.5
86.5
Hook Strength10% weight
90.0
87.7
Length Adherence8% weight
90.4
91.9
Slop Score (EQ-Bench + ours)5% weight
91.9
93.3
Visual Cue Quality4% weight
83.0
88.5

Everything else that differs

GLM-5.3GPT-5.6 Sol (ultra)
Overall / 10088.388.0
Writing Elo22562218
Run-to-run spread (± overall std)1.4801.450
Cost per script (USD)0.1140.363
Avg latency (s)312.8195.5
Open weightsYesNo

Full scorecards: GLM-5.3 · GPT-5.6 Sol (ultra). How scoring works: methodology.

← All comparisons