Towards AITowards AIToneBench

Kimi K3 vs Qwen3.7 Max (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 81.5.

Kimi K3
#8 Elo 2438 · 89.4/100
Qwen3.7 Max (high)
#55 Elo 1674 · 81.5/100
Cost / script
$0.263 vs $0.054
Human baseline
92.7 Kimi K3 falls below it · Qwen3.7 Max (high) falls below it

The verdict

Kimi K3 at #8 against Qwen3.7 Max (high) at #55, 89.39 overall to 81.51. Eight points, intervals nowhere near each other. Anti-slop is the widest gap at 87.36 against 77.78, and continuity is close behind, 88.28 to 79.16. Qwen's best showing is hooks at 84.99, about five and a half behind, so it opens competently and then thins out. Length adherence is respectable at 83.34 but still eight below Kimi's 90.46. On price Qwen is about five cents against Kimi's 26, so five times cheaper, and it's roughly twice as fast at 157 seconds. The catch is that Qwen is closed weights, so if cost is your reason for looking down the board, the genuinely open cheap models like GLM-7 or MiniMax undercut it while scoring in the same range. Qwen's case here is really about already being on Alibaba infrastructure.

Pick Kimi K3 if you want open weights near the top of the board and drafts that hold together across a long script.
Pick Qwen3.7 Max (high) if you're already on Alibaba tooling and want five times cheaper drafts with decent openings.

Metric by metric

Blue bars: Kimi K3. Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.6
80.9
Writing Craft & Clarity13% weight
89.3
81.5
Substance, Accuracy & Value15% weight
89.5
81.2
Continuity & Emotion14% weight
88.3
79.2
YouTube Best Practices12% weight
88.4
80.7
Hook Strength10% weight
89.9
85.0
Length Adherence8% weight
90.5
83.3
Slop Score (EQ-Bench + ours)5% weight
92.0
85.9
Visual Cue Quality4% weight
88.3
78.4

Everything else that differs

Kimi K3Qwen3.7 Max (high)
Overall / 10089.481.5
Writing Elo24381674
Run-to-run spread (± overall std)1.5702.580
Cost per script (USD)0.2630.054
Avg latency (s)367.1157.2
Open weightsYesNo

Full scorecards: Kimi K3 · Qwen3.7 Max (high). How scoring works: methodology.

← All comparisons