Towards AITowards AIToneBench

Kimi K3 vs GLM-5.3

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 88.2 to 86.8.

Kimi K3
#11 Elo 2180 · 88.2/100
GLM-5.3
#24 Elo 2042 · 86.8/100
Cost / script
$0.260 vs $0.109
Human baseline
90.2 Kimi K3 falls below it · GLM-5.3 falls below it

The verdict

Both of these are open weights, so this is a fight inside the open-model column, and Kimi K3 wins it cleanly. It sits at #11 with 2180.2 Elo against GLM-5.3 at #24 with 2041.5, and the confidence intervals don't overlap — Kimi's floor of 2146.3 clears GLM's ceiling of 2083.1 — so the 138.7-point gap is real. Overall it's 88.19 to 86.83, a 1.36-point gap, and Kimi takes nearly every metric: visual cues by the widest margin, 84.6 vs 81.43, length adherence at 90.07 vs 87.5, and YouTube best practices at 87.35 vs 85.02. GLM only answers on the slop side, 91.64 to 90.86 and 96.53 to 95.45 on the slop score. Consistency is the quieter story: Kimi's overall deviation is 1.68 against GLM's 6.6, so GLM swings between good drafts and rough ones while Kimi stays level. Then the bill. GLM runs about eleven cents per script against Kimi's twenty-six, less than half the price. If every draft gets a human pass anyway, that discount is a real argument. If you want the script closer to done on arrival, Kimi is the safer open model.

Pick Kimi K3 if you want the stronger open-weights writer, ahead on nearly every metric and far more consistent draft to draft, at about twenty-six cents per script.
Pick GLM-5.3 if you want open weights at less than half the cost per script, slightly cleaner slop numbers, and can accept output that swings more between drafts.

Metric by metric

Blue bars: Kimi K3. Orange bars: GLM-5.3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
88.3
87.9
Writing Craft & Clarity13% weight
88.5
87.2
Substance, Accuracy & Value15% weight
87.6
85.9
Continuity & Emotion14% weight
86.8
85.1
YouTube Best Practices12% weight
87.3
85.0
Hook Strength10% weight
90.2
89.5
Length Adherence8% weight
90.1
87.5
Slop Score (EQ-Bench + ours)5% weight
90.9
91.6
Visual Cue Quality4% weight
84.6
81.4

Everything else that differs

Kimi K3GLM-5.3
Overall / 10088.286.8
Writing Elo21802042
Run-to-run spread (± overall std)1.6806.600
Cost per script (USD)0.2600.109
Avg latency (s)234.3301.9
Open weightsYesYes

Full scorecards: Kimi K3 · GLM-5.3. How scoring works: methodology.

← All comparisons