Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 88.3 to 88.2. Their confidence intervals overlap, so treat the order as close rather than settled.
This one is closer than the ranks suggest. GLM-5.3 sits at #7 with 2256.0 Elo and 88.34 overall; Kimi K3 is at #17 with 2220.8 Elo and 88.19 overall. Their 95% Elo intervals overlap (2221.7 to 2291.5 against 2177.6 to 2256.4), so do not read much into the exact order. Where GLM-5.3 pulls ahead: Slop Score (EQ-Bench + ours) (91.89 vs 90.86) and Tone & Voice Match (88.72 vs 88.28). Kimi K3 still wins on Visual Cue Quality (84.60 vs 83.02) and Hook Strength (90.19 vs 89.95), so it is not a clean sweep. Price points the same way: GLM-5.3 costs about $0.114 per article against $0.260 for Kimi K3. Both publish open weights. Either is a reasonable default: lean GLM-5.3 for slop score, voice match, and the lower price, Kimi K3 for visual cues and hook strength.
Pick Kimi K3 for visual cues and hook strength.
Pick GLM-5.3 for the stronger board result, slop score, voice match, and the lower price.
Blue bars: GLM-5.3. Orange bars: Kimi K3. Same 0–100 scale; the bold bar wins that metric.
| GLM-5.3 | Kimi K3 | |
|---|---|---|
| Overall / 100 | 88.3 | 88.2 |
| Writing Elo | 2256 | 2221 |
| Run-to-run spread (± overall std) | 1.480 | 1.680 |
| Cost per script (USD) | 0.114 | 0.260 |
| Avg latency (s) | 312.8 | 234.3 |
| Open weights | Yes | Yes |
Full scorecards: GLM-5.3 · Kimi K3. How scoring works: methodology.
← All comparisons