Towards AITowards AIToneBench

Kimi K3 (thinking) vs Qwen3.7 Max (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 81.7.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
Qwen3.7 Max (default)
#50 Elo 1699 · 81.7/100
Cost / script
$0.187 vs $0.054
Human baseline
92.7 Kimi K3 (thinking) falls below it · Qwen3.7 Max (default) falls below it

The verdict

Kimi K3 in thinking mode at #12 against Qwen3.7 Max at #50. 89.25 overall to 81.73, so eight points, and 2416.4 Elo against 1699.1. The single most lopsided metric is length adherence, and it's Kimi's best trick: 91.10 against 87.67. Nearly seven points, and 91.10 is the strongest length control of any model in these comparisons. On a long brief that's the difference between a draft you trim and a draft you restructure. Continuity is the other clear win, 88.28 to 79.03. Qwen's hooks at 84.25 are its most competitive number, about six back. Price runs the other way: Qwen is about five cents against Kimi's 21, so four times cheaper, and it's twice as fast. But Qwen is closed weights, so the cheap-and-open argument that makes GLM-5 or MiniMax interesting doesn't apply here. You're paying less for less, without the hosting upside.

Pick Kimi K3 (thinking) if length discipline and flow across a long script matter, and you want open weights.
Pick Qwen3.7 Max if you're on Alibaba infrastructure and want four times cheaper drafts with solid openings.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
81.2
Writing Craft & Clarity13% weight
89.3
81.7
Substance, Accuracy & Value15% weight
89.2
80.9
Continuity & Emotion14% weight
88.3
79.0
YouTube Best Practices12% weight
87.8
80.8
Hook Strength10% weight
90.2
84.2
Length Adherence8% weight
91.1
85.8
Slop Score (EQ-Bench + ours)5% weight
91.6
86.0
Visual Cue Quality4% weight
87.7
80.2

Everything else that differs

Kimi K3 (thinking)Qwen3.7 Max (default)
Overall / 10089.281.7
Writing Elo24161699
Run-to-run spread (± overall std)1.6503.060
Cost per script (USD)0.1870.054
Avg latency (s)267.2156.8
Open weightsYesNo

Full scorecards: Kimi K3 (thinking) · Qwen3.7 Max (default). How scoring works: methodology.

← All comparisons