Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 80.2.
I'll be honest, I didn't expect Gemini 3.1 Pro at rank 53 on a writing benchmark. I use Gemini daily for image generation and deep research, and it's genuinely strong there. But judged blind against my published scripts, the panel puts it at 80.2 overall versus Kimi K3's 89.3, and the Elo gap, 1577 against 2416, sits far outside both confidence intervals. The two metrics that hurt most: writing craft at 80.8 versus 89.3, so the prose is flatter, and length adherence at 75.9 versus 91.1, so scripts drift off target. Gemini's real advantages are speed, about a minute per script versus Kimi's five and a half, and price, 12 cents against 22. Neither flips the decision for me, because the thing I'd be paying for here is the writing itself. Kimi is also open weights, which only strengthens the case. Gemini stays in my stack for research and images. For scripts in my voice, this one isn't a debate.
Pick Kimi K3 if the writing is the product; it wins this matchup across the board.
Pick Gemini 3.1 Pro if you need a script in about a minute and treat it as a rough starting point.
Blue bars: Kimi K3 (thinking). Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 (thinking) | Gemini 3.1 Pro (default) | |
|---|---|---|
| Overall / 100 | 89.2 | 80.2 |
| Writing Elo | 2416 | 1577 |
| Run-to-run spread (± overall std) | 1.650 | 2.420 |
| Cost per script (USD) | 0.187 | 0.118 |
| Avg latency (s) | 267.2 | 63.8 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 (thinking) · Gemini 3.1 Pro (default). How scoring works: methodology.
← All comparisons