Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.3 to 84.1.
Kimi K3 wins this one cleanly. It sits seventh on the board at 2240.5 Elo; DeepSeek V4 Pro 0813 (max) is #39 at 1847.7, a gap of 392.8 points with no overlap between the confidence intervals. The overall scores are closer than the Elo suggests: 89.26 against 84.09, a 5.17-point spread. The separation comes from craft details. Kimi K3 holds cue quality at 87.41 where DeepSeek lands at 79.38, and it follows length briefs at 88.99 against 81.21. DeepSeek is also less predictable, with a standard deviation of 6.49 to Kimi's 1.52. The counterargument is price. DeepSeek runs about two cents per script; Kimi K3 costs about twenty-seven cents, roughly 12x more. Both ship open weights, so either can be self-hosted. If quality is the brief, Kimi K3 is the pick. If volume is, DeepSeek earns a look.
Pick Kimi K3 if you want a top-of-board writer with tight, repeatable output and can absorb about twenty-seven cents per script.
Pick DeepSeek V4 Pro 0813 (max) if you are producing at volume and about two cents per script matters more than the occasional uneven draft.
Blue bars: Kimi K3. Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | DeepSeek V4 Pro 0813 (max) | |
|---|---|---|
| Overall / 100 | 89.3 | 84.1 |
| Writing Elo | 2240 | 1848 |
| Run-to-run spread (± overall std) | 1.520 | 6.490 |
| Cost per script (USD) | 0.266 | 0.022 |
| Avg latency (s) | 372.1 | 227.3 |
| Open weights | Yes | Yes |
Full scorecards: Kimi K3 · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.
← All comparisons