Towards AITowards AIToneBench

Kimi K3 (thinking) vs DeepSeek V4 Flash 0731

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 84.2.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
DeepSeek V4 Flash 0731
#34 Elo 1955 · 84.2/100
Cost / script
$0.187 vs $0.005
Human baseline
92.7 Kimi K3 (thinking) falls below it · DeepSeek V4 Flash 0731 falls below it

The verdict

Kimi K3 thinking is #12 and the new DeepSeek V4 Flash 0731 is #34, 89.25 against 84.23. Six and a quarter points. The gap is widest on length adherence, 91.10 against 83.57, and YouTube structure, 87.83 against 81.06. DeepSeek's hooks are the bright spot at 87.55, within three of Kimi, so openings land. What it does not do is sustain: continuity 82.02 against 88.28. Both are open weights, so this is a straight quality-versus-price call rather than a licensing one, and the price gap is enormous: half a cent a script against 19 cents. Forty three times cheaper. For six points on a draft you were going to edit anyway, that is a genuinely hard argument to beat at volume. For a script you want close to final, Kimi is worth the money.

Pick Kimi K3 (thinking) when the draft needs to hold its length and flow with minimal editing.
Pick DeepSeek V4 Flash 0731 for volume. Both are open, but this one is forty three times cheaper with comparable hooks.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
85.4
Writing Craft & Clarity13% weight
89.3
85.3
Substance, Accuracy & Value15% weight
89.2
83.5
Continuity & Emotion14% weight
88.3
82.0
YouTube Best Practices12% weight
87.8
81.1
Hook Strength10% weight
90.2
87.5
Length Adherence8% weight
91.1
83.6
Slop Score (EQ-Bench + ours)5% weight
91.6
88.8
Visual Cue Quality4% weight
87.7
82.4

Everything else that differs

Kimi K3 (thinking)DeepSeek V4 Flash 0731
Overall / 10089.284.2
Writing Elo24161955
Run-to-run spread (± overall std)1.6509.240
Cost per script (USD)0.1870.005
Avg latency (s)267.2289.0
Open weightsYesYes

Full scorecards: Kimi K3 (thinking) · DeepSeek V4 Flash 0731. How scoring works: methodology.

← All comparisons