Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 81.7.
Kimi K3 in thinking mode at #12 against Qwen3.7 Max at #50. 89.25 overall to 81.73, so eight points, and 2416.4 Elo against 1699.1. The single most lopsided metric is length adherence, and it's Kimi's best trick: 91.10 against 87.67. Nearly seven points, and 91.10 is the strongest length control of any model in these comparisons. On a long brief that's the difference between a draft you trim and a draft you restructure. Continuity is the other clear win, 88.28 to 79.03. Qwen's hooks at 84.25 are its most competitive number, about six back. Price runs the other way: Qwen is about five cents against Kimi's 21, so four times cheaper, and it's twice as fast. But Qwen is closed weights, so the cheap-and-open argument that makes GLM-5 or MiniMax interesting doesn't apply here. You're paying less for less, without the hosting upside.
Pick Kimi K3 (thinking) if length discipline and flow across a long script matter, and you want open weights.
Pick Qwen3.7 Max if you're on Alibaba infrastructure and want four times cheaper drafts with solid openings.
Blue bars: Kimi K3 (thinking). Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 (thinking) | Qwen3.7 Max (default) | |
|---|---|---|
| Overall / 100 | 89.2 | 81.7 |
| Writing Elo | 2416 | 1699 |
| Run-to-run spread (± overall std) | 1.650 | 3.060 |
| Cost per script (USD) | 0.187 | 0.054 |
| Avg latency (s) | 267.2 | 156.8 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 (thinking) · Qwen3.7 Max (default). How scoring works: methodology.
← All comparisons