Towards AITowards AIToneBench

Kimi K3 vs DeepSeek V4 Pro 0813 (max)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.3 to 84.1.

Kimi K3
#8 Elo 2240 · 89.3/100
DeepSeek V4 Pro 0813 (max)
#38 Elo 1848 · 84.1/100
Cost / script
$0.266 vs $0.022
Human baseline
90.3 Kimi K3 falls below it · DeepSeek V4 Pro 0813 (max) falls below it

The verdict

Kimi K3 wins this one cleanly. It sits seventh on the board at 2240.5 Elo; DeepSeek V4 Pro 0813 (max) is #39 at 1847.7, a gap of 392.8 points with no overlap between the confidence intervals. The overall scores are closer than the Elo suggests: 89.26 against 84.09, a 5.17-point spread. The separation comes from craft details. Kimi K3 holds cue quality at 87.41 where DeepSeek lands at 79.38, and it follows length briefs at 88.99 against 81.21. DeepSeek is also less predictable, with a standard deviation of 6.49 to Kimi's 1.52. The counterargument is price. DeepSeek runs about two cents per script; Kimi K3 costs about twenty-seven cents, roughly 12x more. Both ship open weights, so either can be self-hosted. If quality is the brief, Kimi K3 is the pick. If volume is, DeepSeek earns a look.

Pick Kimi K3 if you want a top-of-board writer with tight, repeatable output and can absorb about twenty-seven cents per script.
Pick DeepSeek V4 Pro 0813 (max) if you are producing at volume and about two cents per script matters more than the occasional uneven draft.

Metric by metric

Blue bars: Kimi K3. Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.5
85.6
Writing Craft & Clarity13% weight
89.4
85.6
Substance, Accuracy & Value15% weight
89.5
83.4
Continuity & Emotion14% weight
88.3
81.8
YouTube Best Practices12% weight
88.7
81.9
Hook Strength10% weight
90.1
88.0
Length Adherence8% weight
89.0
81.2
Slop Score (EQ-Bench + ours)5% weight
91.9
88.8
Visual Cue Quality4% weight
87.4
79.4

Everything else that differs

Kimi K3DeepSeek V4 Pro 0813 (max)
Overall / 10089.384.1
Writing Elo22401848
Run-to-run spread (± overall std)1.5206.490
Cost per script (USD)0.2660.022
Avg latency (s)372.1227.3
Open weightsYesYes

Full scorecards: Kimi K3 · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.

← All comparisons