Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 88.2 to 88.0. Their confidence intervals overlap, so treat the order as close rather than settled.
This one is closer than the ranks suggest. DeepSeek V4.1 Flash (max) sits at #14 with 2226.8 Elo and 88.03 overall; Kimi K3 is at #17 with 2220.8 Elo and 88.19 overall. Their 95% Elo intervals overlap (2185.5 to 2260.8 against 2177.6 to 2256.4), so do not read much into the exact order. Where DeepSeek V4.1 Flash (max) pulls ahead: Length Adherence (91.98 vs 90.07) and Slop Score (EQ-Bench + ours) (91.36 vs 90.86). Kimi K3 still wins on Hook Strength (90.19 vs 89.07) and Visual Cue Quality (84.60 vs 83.68), so it is not a clean sweep. Price points the same way: DeepSeek V4.1 Flash (max) costs about $0.013 per article against $0.260 for Kimi K3. Both publish open weights. Either is a reasonable default: lean DeepSeek V4.1 Flash (max) for length adherence, slop score, and the lower price, Kimi K3 for hook strength and visual cues.
Pick Kimi K3 for hook strength and visual cues.
Pick DeepSeek V4.1 Flash (max) for the stronger board result, length adherence, slop score, and the lower price.
Blue bars: DeepSeek V4.1 Flash (max). Orange bars: Kimi K3. Same 0–100 scale; the bold bar wins that metric.
| DeepSeek V4.1 Flash (max) | Kimi K3 | |
|---|---|---|
| Overall / 100 | 88.0 | 88.2 |
| Writing Elo | 2227 | 2221 |
| Run-to-run spread (± overall std) | 1.280 | 1.680 |
| Cost per script (USD) | 0.013 | 0.260 |
| Avg latency (s) | 76.2 | 234.3 |
| Open weights | Yes | Yes |
Full scorecards: DeepSeek V4.1 Flash (max) · Kimi K3. How scoring works: methodology.
← All comparisons