Towards AITowards AIToneBench

Kimi K3 (thinking) vs DeepSeek V4 Pro (xhigh)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 80.5.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
DeepSeek V4 Pro (xhigh)
#60 Elo 1601 · 80.5/100
Cost / script
$0.187 vs $0.014
Human baseline
92.7 Kimi K3 (thinking) falls below it · DeepSeek V4 Pro (xhigh) falls below it

The verdict

Kimi K3 is the best open-weights model on my board, and the only open one near the top. With thinking on it lands seventh, and all six models above it are closed Anthropic configs. The lock at the top holds, Kimi just gets closest to it. The Elo gap over DeepSeek V4 Pro sits well outside both confidence intervals. Reading them back to back, two things stand out. Kimi respects the length target, 91.10 on length adherence against 80.47, which matters because overlong scripts are the first thing I have to cut. And its scripts flow like one story instead of stacked sections, 88.28 versus 78.51 on continuity and emotion, the thing I find hardest to fix in an edit. DeepSeek's defense is price: about a cent and a half per script versus 19 cents, and it finishes faster since Kimi spends a long time thinking. Fifteen times the cost sounds dramatic until you remember both are cheap in absolute terms. For anything I would publish, I take Kimi. For bulk drafting with an editor downstream, DeepSeek earns its slot.

Pick Kimi K3 (thinking) if you want the strongest open-weights writing available and can live with the slower thinking time.
Pick DeepSeek V4 Pro (xhigh) if you draft in bulk and a cent and a half per script matters more than polish.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
82.2
Writing Craft & Clarity13% weight
89.3
82.2
Substance, Accuracy & Value15% weight
89.2
81.0
Continuity & Emotion14% weight
88.3
78.5
YouTube Best Practices12% weight
87.8
76.9
Hook Strength10% weight
90.2
84.6
Length Adherence8% weight
91.1
80.5
Slop Score (EQ-Bench + ours)5% weight
91.6
79.4
Visual Cue Quality4% weight
87.7
73.6

Everything else that differs

Kimi K3 (thinking)DeepSeek V4 Pro (xhigh)
Overall / 10089.280.5
Writing Elo24161601
Run-to-run spread (± overall std)1.6503.670
Cost per script (USD)0.1870.014
Avg latency (s)267.2163.2
Open weightsYesYes

Full scorecards: Kimi K3 (thinking) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.

← All comparisons