Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.4 to 82.4.
GPT-5.6 Sol at ultra is #16, GLM-5 is #43, and the overall is 88.40 against 82.42. Six points. The single number that should decide this is anti-slop: 90.60 against 75.66. Fourteen points. GLM writes noticeably more AI-flavoured filler and is less careful with figures, and that is the kind of thing an editor fixes line by line rather than in one pass. Length adherence is the other collapse, 94.34 against 78.61. GLM's hooks are decent at 86.47, only slightly behind, so openings usually survive. Cost runs hard the other way: just over two cents a script against GPT's 16, so six times cheaper, plus open weights and much faster at 117 seconds against 421. If your pipeline has real editorial capacity and volume is the constraint, GLM at two cents is a serious option. If the draft needs to be close to shippable, the slop gap costs more in editing time than you saved.
Pick GPT-5.6 Sol (ultra) if you want drafts that are close to clean, especially on slop, numbers and length.
Pick GLM-5 if cost and speed dominate and you have an editor. Open weights, six times cheaper, three times faster.
Blue bars: GPT-5.6 Sol (ultra). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (ultra) | GLM-5 | |
|---|---|---|
| Overall / 100 | 88.4 | 82.4 |
| Writing Elo | 2316 | 1783 |
| Run-to-run spread (± overall std) | 1.970 | 3.760 |
| Cost per script (USD) | 0.137 | 0.023 |
| Avg latency (s) | 420.9 | 117.3 |
| Open weights | No | Yes |
Full scorecards: GPT-5.6 Sol (ultra) · GLM-5. How scoring works: methodology.
← All comparisons