Towards AITowards AIToneBench

Kimi K3 vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 83.4.

Kimi K3
#8 Elo 2438 · 89.4/100
Qwen3.8 Max
#38 Elo 1866 · 83.4/100
Cost / script
$0.263 vs $0.150
Human baseline
92.7 Kimi K3 falls below it · Qwen3.8 Max falls below it

The verdict

Kimi K3 at #8 with 89.39 overall, Qwen3.8 Max at #38 with 83.44, and six points at this end of the board is not a close call. Kimi wins everything that reads as writing: hooks 89.88 against 83.50, tone 89.61 against 81.93, and continuity 88.28 against 78.08, which is a ten-point gap on the metric that decides whether a script holds attention or wanders. Qwen's answers are discipline metrics: it edges number handling 87.78 to 87.36 and length adherence 92.90 to 90.46, so its drafts arrive the right size with clean stats. That profile has a real use: templated scripts where structure is imposed from outside. But it also costs 15 cents a script against Kimi's 26, which is the wrong price for a six-point gap. Where the budget allows either, this one is simple. Kimi writes, Qwen formats.

Pick Kimi K3 for anything a viewer will actually watch: six points better overall and a ten-point edge on continuity.
Pick Qwen3.8 Max only for tightly templated output: excellent length control and clean numbers at a bit over half Kimi's price.

Metric by metric

Blue bars: Kimi K3. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.6
81.9
Writing Craft & Clarity13% weight
89.3
83.4
Substance, Accuracy & Value15% weight
89.5
84.4
Continuity & Emotion14% weight
88.3
78.1
YouTube Best Practices12% weight
88.4
80.4
Hook Strength10% weight
89.9
83.5
Length Adherence8% weight
90.5
92.9
Slop Score (EQ-Bench + ours)5% weight
92.0
92.0
Visual Cue Quality4% weight
88.3
85.2

Everything else that differs

Kimi K3Qwen3.8 Max
Overall / 10089.483.4
Writing Elo24381866
Run-to-run spread (± overall std)1.5705.620
Cost per script (USD)0.2630.150
Avg latency (s)367.1397.4
Open weightsYesNo

Full scorecards: Kimi K3 · Qwen3.8 Max. How scoring works: methodology.

← All comparisons