Towards AITowards AIToneBench

Kimi K3 vs Gemini 3.1 Pro (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 80.2.

Kimi K3
#8 Elo 2438 · 89.4/100
Gemini 3.1 Pro (default)
#63 Elo 1577 · 80.2/100
Cost / script
$0.263 vs $0.118
Human baseline
92.7 Kimi K3 falls below it · Gemini 3.1 Pro (default) falls below it

The verdict

Kimi K3 is #8, Gemini 3.1 Pro is #63, and the spread is nine points of overall score, 89.39 against 80.21. That reads harsher than most people expect from Gemini, which is a strong model on plenty of other work. On this job the failure is specific: length adherence at 75.91 against Kimi's 90.46, and continuity at 78.40 against 88.28. Fourteen and ten points. It writes fine paragraphs and does not hold a long script together or land near a target length. Its cue quality at 81.66 is the part that holds up best. Gemini is much faster, 60 seconds against 367, and less than half the price, 12 cents against 26. So it's the quick cheap draft against the one that needs less work afterward. On a 4,000 word brief I know which side of that I'd rather be on, but if you're generating a lot of short material the calculus changes.

Pick Kimi K3 for long-form, where holding length and flow across the whole script is the actual job.
Pick Gemini 3.1 Pro if you want drafts back in about a minute at half the cost and you're editing structure anyway.

Metric by metric

Blue bars: Kimi K3. Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.6
79.9
Writing Craft & Clarity13% weight
89.3
80.8
Substance, Accuracy & Value15% weight
89.5
80.8
Continuity & Emotion14% weight
88.3
78.4
YouTube Best Practices12% weight
88.4
78.9
Hook Strength10% weight
89.9
82.3
Length Adherence8% weight
90.5
75.9
Slop Score (EQ-Bench + ours)5% weight
92.0
87.7
Visual Cue Quality4% weight
88.3
81.7

Everything else that differs

Kimi K3Gemini 3.1 Pro (default)
Overall / 10089.480.2
Writing Elo24381577
Run-to-run spread (± overall std)1.5702.420
Cost per script (USD)0.2630.118
Avg latency (s)367.163.8
Open weightsYesNo

Full scorecards: Kimi K3 · Gemini 3.1 Pro (default). How scoring works: methodology.

← All comparisons