Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 83.4.
Kimi K3 in thinking mode at #12 against Qwen3.8 Max at #38, 89.25 overall to 83.44. Just under six points. Qwen's calling card is length, and it genuinely wins there: 92.90 against 91.10, which is a real achievement given Kimi's length control is one of the best on the board. Everything else goes to Kimi, and the decisive metric is continuity, 88.28 against 78.08. Over ten points. Voice is the other big one, 89.31 against 81.93. So Qwen produces correctly-sized scripts that read as assembled rather than spoken. Practically: Qwen costs 15 cents against Kimi's 21, so a third cheaper, and they are similar on speed. But Kimi is open weights and Qwen is not, which flips the usual reason you would go down-board for cost.
Pick Kimi K3 (thinking) if flow and voice matter, and you want open weights near the top of the board.
Pick Qwen3.8 Max for the best length adherence of the two at a third less cost.
Blue bars: Kimi K3 (thinking). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 (thinking) | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 89.2 | 83.4 |
| Writing Elo | 2416 | 1866 |
| Run-to-run spread (± overall std) | 1.650 | 5.620 |
| Cost per script (USD) | 0.187 | 0.150 |
| Avg latency (s) | 267.2 | 397.4 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 (thinking) · Qwen3.8 Max. How scoring works: methodology.
← All comparisons