Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 81.5.
Kimi K3 at #8 against Qwen3.7 Max (high) at #55, 89.39 overall to 81.51. Eight points, intervals nowhere near each other. Anti-slop is the widest gap at 87.36 against 77.78, and continuity is close behind, 88.28 to 79.16. Qwen's best showing is hooks at 84.99, about five and a half behind, so it opens competently and then thins out. Length adherence is respectable at 83.34 but still eight below Kimi's 90.46. On price Qwen is about five cents against Kimi's 26, so five times cheaper, and it's roughly twice as fast at 157 seconds. The catch is that Qwen is closed weights, so if cost is your reason for looking down the board, the genuinely open cheap models like GLM-7 or MiniMax undercut it while scoring in the same range. Qwen's case here is really about already being on Alibaba infrastructure.
Pick Kimi K3 if you want open weights near the top of the board and drafts that hold together across a long script.
Pick Qwen3.7 Max (high) if you're already on Alibaba tooling and want five times cheaper drafts with decent openings.
Blue bars: Kimi K3. Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | Qwen3.7 Max (high) | |
|---|---|---|
| Overall / 100 | 89.4 | 81.5 |
| Writing Elo | 2438 | 1674 |
| Run-to-run spread (± overall std) | 1.570 | 2.580 |
| Cost per script (USD) | 0.263 | 0.054 |
| Avg latency (s) | 367.1 | 157.2 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 · Qwen3.7 Max (high). How scoring works: methodology.
← All comparisons