Towards AITowards AIToneBench

Kimi K3 vs DeepSeek V4 Flash 0731

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 84.2.

Kimi K3
#8 Elo 2438 · 89.4/100
DeepSeek V4 Flash 0731
#34 Elo 1955 · 84.2/100
Cost / script
$0.263 vs $0.005
Human baseline
92.7 Kimi K3 falls below it · DeepSeek V4 Flash 0731 falls below it

The verdict

The best open-weights model on the board against the cheapest one worth using, and the gap is honest: Kimi K3 at #8 with 89.39 overall, DeepSeek V4 Flash 0731 at #34 with 84.23, five points, intervals nowhere near each other. Kimi is simply the better writer everywhere I look: tone 89.61 against 85.44, continuity 88.28 against 82.02, YouTube structure 88.39 against 81.06. DeepSeek's hooks at 87.55 are the one place it plays in Kimi's league. So why is this even a page? Price. DeepSeek runs about half a cent a script against Kimi's 26 cents, roughly fifty times cheaper, and the 0731 build just climbed 49 places in one release. Five points is a real quality gap; fifty times is a real budget gap. If the script ships to my channel, I pay for Kimi. If I'm bulk-drafting and a human owns the rewrite anyway, DeepSeek's math is unbeatable.

Pick Kimi K3 when the draft has to stand on its own: it wins every writing metric by a wide margin and is still open weights.
Pick DeepSeek V4 Flash 0731 for volume drafting with a human editor: about fifty times cheaper, respectable hooks, and improving fast.

Metric by metric

Blue bars: Kimi K3. Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.6
85.4
Writing Craft & Clarity13% weight
89.3
85.3
Substance, Accuracy & Value15% weight
89.5
83.5
Continuity & Emotion14% weight
88.3
82.0
YouTube Best Practices12% weight
88.4
81.1
Hook Strength10% weight
89.9
87.5
Length Adherence8% weight
90.5
83.6
Slop Score (EQ-Bench + ours)5% weight
92.0
88.8
Visual Cue Quality4% weight
88.3
82.4

Everything else that differs

Kimi K3DeepSeek V4 Flash 0731
Overall / 10089.484.2
Writing Elo24381955
Run-to-run spread (± overall std)1.5709.240
Cost per script (USD)0.2630.005
Avg latency (s)367.1289.0
Open weightsYesYes

Full scorecards: Kimi K3 · DeepSeek V4 Flash 0731. How scoring works: methodology.

← All comparisons