Towards AITowards AIToneBench

Kimi K3 (thinking) vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 82.4.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
GLM-5
#43 Elo 1783 · 82.4/100
Cost / script
$0.187 vs $0.023
Human baseline
92.7 Kimi K3 (thinking) falls below it · GLM-5 falls below it

The verdict

Both of these are open weights, which I love seeing, but the writing gap is real. Kimi K3 scores 89.2 overall; GLM-5 sits at 82.4, and the Elo intervals are nowhere near each other, so no uncertainty to hide behind. GLM's most telling number is 77.3 on number handling and slop, with one of the widest spreads of any metric in this matchup: some scripts recite stats like a quarterly report, others behave fine, and you don't know which you'll get. Continuity and emotion show the same softness, 79.7 versus Kimi's 88.3, facts in a row instead of a story. Now, the case for GLM is price. At under 2 cents per script, an estimate to be fair, it's roughly thirteen times cheaper than Kimi, and it still opens well, 86.5 on hooks. For drafts you'll rewrite yourself anyway, that math genuinely works. For anything close to final copy in my voice, Kimi is the only pick of the two.

Pick Kimi K3 if you want the strongest open-weights writer here with consistent, publishable structure.
Pick GLM-5 if budget rules the decision and every script gets a human rewrite anyway.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
83.1
Writing Craft & Clarity13% weight
89.3
84.0
Substance, Accuracy & Value15% weight
89.2
84.0
Continuity & Emotion14% weight
88.3
79.7
YouTube Best Practices12% weight
87.8
78.3
Hook Strength10% weight
90.2
86.5
Length Adherence8% weight
91.1
78.6
Slop Score (EQ-Bench + ours)5% weight
91.6
85.3
Visual Cue Quality4% weight
87.7
83.8

Everything else that differs

Kimi K3 (thinking)GLM-5
Overall / 10089.282.4
Writing Elo24161783
Run-to-run spread (± overall std)1.6503.760
Cost per script (USD)0.1870.023
Avg latency (s)267.2117.3
Open weightsYesYes

Full scorecards: Kimi K3 (thinking) · GLM-5. How scoring works: methodology.

← All comparisons