Towards AITowards AIToneBench

Kimi K3 vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 82.4.

Kimi K3
#8 Elo 2438 · 89.4/100
GLM-5
#43 Elo 1783 · 82.4/100
Cost / script
$0.263 vs $0.023
Human baseline
92.7 Kimi K3 falls below it · GLM-5 falls below it

The verdict

Two open-weights models, very different price points. Kimi K3 is #8 at 89.39 overall, GLM-5 is #43 at 82.42, and GLM costs about two cents a script against Kimi's 25. So you're asking whether seven points of quality is worth roughly eleven times the price. The gap concentrates in one place: anti-slop and numbers, 87.36 against 75.66. Eleven and a half points. GLM writes noticeably more AI-flavoured filler and is less reliable with figures. Its YouTube structure score of 78.28 against Kimi's 88.39 is the other big one. What GLM does hold is hooks at 86.47, only three and a half behind, so openings are usually fine. GLM is also three times faster, 117 seconds against 367. For a pipeline with real editorial capacity, GLM at two cents is a serious option. For drafts that need to be close to done, Kimi is worth the eleven times.

Pick Kimi K3 if you want open weights that draft close to shippable, especially on slop and structure.
Pick GLM-5 if you're running volume on a budget. Two cents a script, three times faster, and hooks that hold up.

Metric by metric

Blue bars: Kimi K3. Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.6
83.1
Writing Craft & Clarity13% weight
89.3
84.0
Substance, Accuracy & Value15% weight
89.5
84.0
Continuity & Emotion14% weight
88.3
79.7
YouTube Best Practices12% weight
88.4
78.3
Hook Strength10% weight
89.9
86.5
Length Adherence8% weight
90.5
78.6
Slop Score (EQ-Bench + ours)5% weight
92.0
85.3
Visual Cue Quality4% weight
88.3
83.8

Everything else that differs

Kimi K3GLM-5
Overall / 10089.482.4
Writing Elo24381783
Run-to-run spread (± overall std)1.5703.760
Cost per script (USD)0.2630.023
Avg latency (s)367.1117.3
Open weightsYesYes

Full scorecards: Kimi K3 · GLM-5. How scoring works: methodology.

← All comparisons