Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 83.4.
Kimi K3 at #8 with 89.39 overall, Qwen3.8 Max at #38 with 83.44, and six points at this end of the board is not a close call. Kimi wins everything that reads as writing: hooks 89.88 against 83.50, tone 89.61 against 81.93, and continuity 88.28 against 78.08, which is a ten-point gap on the metric that decides whether a script holds attention or wanders. Qwen's answers are discipline metrics: it edges number handling 87.78 to 87.36 and length adherence 92.90 to 90.46, so its drafts arrive the right size with clean stats. That profile has a real use: templated scripts where structure is imposed from outside. But it also costs 15 cents a script against Kimi's 26, which is the wrong price for a six-point gap. Where the budget allows either, this one is simple. Kimi writes, Qwen formats.
Pick Kimi K3 for anything a viewer will actually watch: six points better overall and a ten-point edge on continuity.
Pick Qwen3.8 Max only for tightly templated output: excellent length control and clean numbers at a bit over half Kimi's price.
Blue bars: Kimi K3. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 89.4 | 83.4 |
| Writing Elo | 2438 | 1866 |
| Run-to-run spread (± overall std) | 1.570 | 5.620 |
| Cost per script (USD) | 0.263 | 0.150 |
| Avg latency (s) | 367.1 | 397.4 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 · Qwen3.8 Max. How scoring works: methodology.
← All comparisons