Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 81.5.
This is one of the widest voice gaps in this set. Kimi K3 scores 89.0 on tone against Qwen3.7 Max's 80.7, eight full points, and tone carries the most weight on ToneBench because sounding like me is the whole assignment. Writing craft follows the same pattern, 89.3 versus 81.5. The shape of Qwen's scores reads as competent but generic: it hooks reasonably well at 85.0 and holds its target length, so the skeleton of a script is there. The person in it isn't. And the Elo intervals sit nowhere near each other, so I don't need to hedge on the ranking. The odd part of this matchup is that the usual trade-off is missing: Kimi is open weights and Qwen isn't, so going with Qwen doesn't even buy you openness. What it buys you is price, about 5 cents per script, estimated, versus Kimi's 22. If your pipeline rewrites the voice layer anyway, that can work. Otherwise I don't really see the case.
Pick Kimi K3 if voice matters at all; it leads on nearly every metric and is open weights on top.
Pick Qwen3.7 Max if you only need a cheap structural draft and your pipeline handles the voice layer.
Blue bars: Kimi K3 (thinking). Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 (thinking) | Qwen3.7 Max (high) | |
|---|---|---|
| Overall / 100 | 89.2 | 81.5 |
| Writing Elo | 2416 | 1674 |
| Run-to-run spread (± overall std) | 1.650 | 2.580 |
| Cost per script (USD) | 0.187 | 0.054 |
| Avg latency (s) | 267.2 | 157.2 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 (thinking) · Qwen3.7 Max (high). How scoring works: methodology.
← All comparisons