Towards AITowards AIToneBench

Kimi K3 (thinking) vs Qwen3.7 Max (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 81.5.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
Qwen3.7 Max (high)
#55 Elo 1674 · 81.5/100
Cost / script
$0.187 vs $0.054
Human baseline
92.7 Kimi K3 (thinking) falls below it · Qwen3.7 Max (high) falls below it

The verdict

This is one of the widest voice gaps in this set. Kimi K3 scores 89.0 on tone against Qwen3.7 Max's 80.7, eight full points, and tone carries the most weight on ToneBench because sounding like me is the whole assignment. Writing craft follows the same pattern, 89.3 versus 81.5. The shape of Qwen's scores reads as competent but generic: it hooks reasonably well at 85.0 and holds its target length, so the skeleton of a script is there. The person in it isn't. And the Elo intervals sit nowhere near each other, so I don't need to hedge on the ranking. The odd part of this matchup is that the usual trade-off is missing: Kimi is open weights and Qwen isn't, so going with Qwen doesn't even buy you openness. What it buys you is price, about 5 cents per script, estimated, versus Kimi's 22. If your pipeline rewrites the voice layer anyway, that can work. Otherwise I don't really see the case.

Pick Kimi K3 if voice matters at all; it leads on nearly every metric and is open weights on top.
Pick Qwen3.7 Max if you only need a cheap structural draft and your pipeline handles the voice layer.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
80.9
Writing Craft & Clarity13% weight
89.3
81.5
Substance, Accuracy & Value15% weight
89.2
81.2
Continuity & Emotion14% weight
88.3
79.2
YouTube Best Practices12% weight
87.8
80.7
Hook Strength10% weight
90.2
85.0
Length Adherence8% weight
91.1
83.3
Slop Score (EQ-Bench + ours)5% weight
91.6
85.9
Visual Cue Quality4% weight
87.7
78.4

Everything else that differs

Kimi K3 (thinking)Qwen3.7 Max (high)
Overall / 10089.281.5
Writing Elo24161674
Run-to-run spread (± overall std)1.6502.580
Cost per script (USD)0.1870.054
Avg latency (s)267.2157.2
Open weightsYesNo

Full scorecards: Kimi K3 (thinking) · Qwen3.7 Max (high). How scoring works: methodology.

← All comparisons