Towards AITowards AIToneBench

Kimi K3 (thinking) vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 83.4.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
Qwen3.8 Max
#38 Elo 1866 · 83.4/100
Cost / script
$0.187 vs $0.150
Human baseline
92.7 Kimi K3 (thinking) falls below it · Qwen3.8 Max falls below it

The verdict

Kimi K3 in thinking mode at #12 against Qwen3.8 Max at #38, 89.25 overall to 83.44. Just under six points. Qwen's calling card is length, and it genuinely wins there: 92.90 against 91.10, which is a real achievement given Kimi's length control is one of the best on the board. Everything else goes to Kimi, and the decisive metric is continuity, 88.28 against 78.08. Over ten points. Voice is the other big one, 89.31 against 81.93. So Qwen produces correctly-sized scripts that read as assembled rather than spoken. Practically: Qwen costs 15 cents against Kimi's 21, so a third cheaper, and they are similar on speed. But Kimi is open weights and Qwen is not, which flips the usual reason you would go down-board for cost.

Pick Kimi K3 (thinking) if flow and voice matter, and you want open weights near the top of the board.
Pick Qwen3.8 Max for the best length adherence of the two at a third less cost.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
81.9
Writing Craft & Clarity13% weight
89.3
83.4
Substance, Accuracy & Value15% weight
89.2
84.4
Continuity & Emotion14% weight
88.3
78.1
YouTube Best Practices12% weight
87.8
80.4
Hook Strength10% weight
90.2
83.5
Length Adherence8% weight
91.1
92.9
Slop Score (EQ-Bench + ours)5% weight
91.6
92.0
Visual Cue Quality4% weight
87.7
85.2

Everything else that differs

Kimi K3 (thinking)Qwen3.8 Max
Overall / 10089.283.4
Writing Elo24161866
Run-to-run spread (± overall std)1.6505.620
Cost per script (USD)0.1870.150
Avg latency (s)267.2397.4
Open weightsYesNo

Full scorecards: Kimi K3 (thinking) · Qwen3.8 Max. How scoring works: methodology.

← All comparisons