Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 88.2 to 86.8.
Both of these are open weights, so this is a fight inside the open-model column, and Kimi K3 wins it cleanly. It sits at #11 with 2180.2 Elo against GLM-5.3 at #24 with 2041.5, and the confidence intervals don't overlap — Kimi's floor of 2146.3 clears GLM's ceiling of 2083.1 — so the 138.7-point gap is real. Overall it's 88.19 to 86.83, a 1.36-point gap, and Kimi takes nearly every metric: visual cues by the widest margin, 84.6 vs 81.43, length adherence at 90.07 vs 87.5, and YouTube best practices at 87.35 vs 85.02. GLM only answers on the slop side, 91.64 to 90.86 and 96.53 to 95.45 on the slop score. Consistency is the quieter story: Kimi's overall deviation is 1.68 against GLM's 6.6, so GLM swings between good drafts and rough ones while Kimi stays level. Then the bill. GLM runs about eleven cents per script against Kimi's twenty-six, less than half the price. If every draft gets a human pass anyway, that discount is a real argument. If you want the script closer to done on arrival, Kimi is the safer open model.
Pick Kimi K3 if you want the stronger open-weights writer, ahead on nearly every metric and far more consistent draft to draft, at about twenty-six cents per script.
Pick GLM-5.3 if you want open weights at less than half the cost per script, slightly cleaner slop numbers, and can accept output that swings more between drafts.
Blue bars: Kimi K3. Orange bars: GLM-5.3. Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | GLM-5.3 | |
|---|---|---|
| Overall / 100 | 88.2 | 86.8 |
| Writing Elo | 2180 | 2042 |
| Run-to-run spread (± overall std) | 1.680 | 6.600 |
| Cost per script (USD) | 0.260 | 0.109 |
| Avg latency (s) | 234.3 | 301.9 |
| Open weights | Yes | Yes |
Full scorecards: Kimi K3 · GLM-5.3. How scoring works: methodology.
← All comparisons