Towards AITowards AIToneBench

Kimi K3 (thinking) vs Gemini 3.1 Pro (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 80.2.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
Gemini 3.1 Pro (default)
#63 Elo 1577 · 80.2/100
Cost / script
$0.187 vs $0.118
Human baseline
92.7 Kimi K3 (thinking) falls below it · Gemini 3.1 Pro (default) falls below it

The verdict

I'll be honest, I didn't expect Gemini 3.1 Pro at rank 53 on a writing benchmark. I use Gemini daily for image generation and deep research, and it's genuinely strong there. But judged blind against my published scripts, the panel puts it at 80.2 overall versus Kimi K3's 89.3, and the Elo gap, 1577 against 2416, sits far outside both confidence intervals. The two metrics that hurt most: writing craft at 80.8 versus 89.3, so the prose is flatter, and length adherence at 75.9 versus 91.1, so scripts drift off target. Gemini's real advantages are speed, about a minute per script versus Kimi's five and a half, and price, 12 cents against 22. Neither flips the decision for me, because the thing I'd be paying for here is the writing itself. Kimi is also open weights, which only strengthens the case. Gemini stays in my stack for research and images. For scripts in my voice, this one isn't a debate.

Pick Kimi K3 if the writing is the product; it wins this matchup across the board.
Pick Gemini 3.1 Pro if you need a script in about a minute and treat it as a rough starting point.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
79.9
Writing Craft & Clarity13% weight
89.3
80.8
Substance, Accuracy & Value15% weight
89.2
80.8
Continuity & Emotion14% weight
88.3
78.4
YouTube Best Practices12% weight
87.8
78.9
Hook Strength10% weight
90.2
82.3
Length Adherence8% weight
91.1
75.9
Slop Score (EQ-Bench + ours)5% weight
91.6
87.7
Visual Cue Quality4% weight
87.7
81.7

Everything else that differs

Kimi K3 (thinking)Gemini 3.1 Pro (default)
Overall / 10089.280.2
Writing Elo24161577
Run-to-run spread (± overall std)1.6502.420
Cost per script (USD)0.1870.118
Avg latency (s)267.263.8
Open weightsYesNo

Full scorecards: Kimi K3 (thinking) · Gemini 3.1 Pro (default). How scoring works: methodology.

← All comparisons