Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 88.3 to 88.0. Their confidence intervals overlap, so treat the order as close rather than settled.
This one is closer than the ranks suggest. GLM-5.3 sits at #7 with 2256.0 Elo and 88.34 overall; GPT-5.6 Sol (ultra) is at #19 with 2218.2 Elo and 87.97 overall. Their 95% Elo intervals overlap (2221.7 to 2291.5 against 2177.5 to 2256.6), so do not read much into the exact order. Where GLM-5.3 pulls ahead: Hook Strength (89.95 vs 87.69) and Tone & Voice Match (88.72 vs 86.80). GPT-5.6 Sol (ultra) still wins on Visual Cue Quality (88.55 vs 83.02) and Length Adherence (91.90 vs 90.43), so it is not a clean sweep. Price points the same way: GLM-5.3 costs about $0.114 per article against $0.363 for GPT-5.6 Sol (ultra). GLM-5.3 publishes open weights; GPT-5.6 Sol (ultra) does not. Either is a reasonable default: lean GLM-5.3 for hook strength, voice match, the lower price, and open weights, GPT-5.6 Sol (ultra) for visual cues and length adherence.
Pick GPT-5.6 Sol (ultra) for visual cues and length adherence.
Pick GLM-5.3 for the stronger board result, hook strength, voice match, open weights, and the lower price.
Blue bars: GLM-5.3. Orange bars: GPT-5.6 Sol (ultra). Same 0–100 scale; the bold bar wins that metric.
| GLM-5.3 | GPT-5.6 Sol (ultra) | |
|---|---|---|
| Overall / 100 | 88.3 | 88.0 |
| Writing Elo | 2256 | 2218 |
| Run-to-run spread (± overall std) | 1.480 | 1.450 |
| Cost per script (USD) | 0.114 | 0.363 |
| Avg latency (s) | 312.8 | 195.5 |
| Open weights | Yes | No |
Full scorecards: GLM-5.3 · GPT-5.6 Sol (ultra). How scoring works: methodology.
← All comparisons