Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 80.5.
Kimi K3 is the best open-weights model on my board, and the only open one near the top. With thinking on it lands seventh, and all six models above it are closed Anthropic configs. The lock at the top holds, Kimi just gets closest to it. The Elo gap over DeepSeek V4 Pro sits well outside both confidence intervals. Reading them back to back, two things stand out. Kimi respects the length target, 91.10 on length adherence against 80.47, which matters because overlong scripts are the first thing I have to cut. And its scripts flow like one story instead of stacked sections, 88.28 versus 78.51 on continuity and emotion, the thing I find hardest to fix in an edit. DeepSeek's defense is price: about a cent and a half per script versus 19 cents, and it finishes faster since Kimi spends a long time thinking. Fifteen times the cost sounds dramatic until you remember both are cheap in absolute terms. For anything I would publish, I take Kimi. For bulk drafting with an editor downstream, DeepSeek earns its slot.
Pick Kimi K3 (thinking) if you want the strongest open-weights writing available and can live with the slower thinking time.
Pick DeepSeek V4 Pro (xhigh) if you draft in bulk and a cent and a half per script matters more than polish.
Blue bars: Kimi K3 (thinking). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 (thinking) | DeepSeek V4 Pro (xhigh) | |
|---|---|---|
| Overall / 100 | 89.2 | 80.5 |
| Writing Elo | 2416 | 1601 |
| Run-to-run spread (± overall std) | 1.650 | 3.670 |
| Cost per script (USD) | 0.187 | 0.014 |
| Avg latency (s) | 267.2 | 163.2 |
| Open weights | Yes | Yes |
Full scorecards: Kimi K3 (thinking) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.
← All comparisons