Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 82.4.
Both of these are open weights, which I love seeing, but the writing gap is real. Kimi K3 scores 89.2 overall; GLM-5 sits at 82.4, and the Elo intervals are nowhere near each other, so no uncertainty to hide behind. GLM's most telling number is 77.3 on number handling and slop, with one of the widest spreads of any metric in this matchup: some scripts recite stats like a quarterly report, others behave fine, and you don't know which you'll get. Continuity and emotion show the same softness, 79.7 versus Kimi's 88.3, facts in a row instead of a story. Now, the case for GLM is price. At under 2 cents per script, an estimate to be fair, it's roughly thirteen times cheaper than Kimi, and it still opens well, 86.5 on hooks. For drafts you'll rewrite yourself anyway, that math genuinely works. For anything close to final copy in my voice, Kimi is the only pick of the two.
Pick Kimi K3 if you want the strongest open-weights writer here with consistent, publishable structure.
Pick GLM-5 if budget rules the decision and every script gets a human rewrite anyway.
Blue bars: Kimi K3 (thinking). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 (thinking) | GLM-5 | |
|---|---|---|
| Overall / 100 | 89.2 | 82.4 |
| Writing Elo | 2416 | 1783 |
| Run-to-run spread (± overall std) | 1.650 | 3.760 |
| Cost per script (USD) | 0.187 | 0.023 |
| Avg latency (s) | 267.2 | 117.3 |
| Open weights | Yes | Yes |
Full scorecards: Kimi K3 (thinking) · GLM-5. How scoring works: methodology.
← All comparisons