Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 82.4.
Two open-weights models, very different price points. Kimi K3 is #8 at 89.39 overall, GLM-5 is #43 at 82.42, and GLM costs about two cents a script against Kimi's 25. So you're asking whether seven points of quality is worth roughly eleven times the price. The gap concentrates in one place: anti-slop and numbers, 87.36 against 75.66. Eleven and a half points. GLM writes noticeably more AI-flavoured filler and is less reliable with figures. Its YouTube structure score of 78.28 against Kimi's 88.39 is the other big one. What GLM does hold is hooks at 86.47, only three and a half behind, so openings are usually fine. GLM is also three times faster, 117 seconds against 367. For a pipeline with real editorial capacity, GLM at two cents is a serious option. For drafts that need to be close to done, Kimi is worth the eleven times.
Pick Kimi K3 if you want open weights that draft close to shippable, especially on slop and structure.
Pick GLM-5 if you're running volume on a budget. Two cents a script, three times faster, and hooks that hold up.
Blue bars: Kimi K3. Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | GLM-5 | |
|---|---|---|
| Overall / 100 | 89.4 | 82.4 |
| Writing Elo | 2438 | 1783 |
| Run-to-run spread (± overall std) | 1.570 | 3.760 |
| Cost per script (USD) | 0.263 | 0.023 |
| Avg latency (s) | 367.1 | 117.3 |
| Open weights | Yes | Yes |
Full scorecards: Kimi K3 · GLM-5. How scoring works: methodology.
← All comparisons