Towards AITowards AIToneBench

Thinking levels · Google

Does Gemini 3.8 Flash write better when it thinks harder?

A little, maybe. High thinking puts Gemini 3.8 Flash 24 Elo ahead of its default, but you pay 1.9 times as much per script for it, and the gap is small enough that it could be noise. Its real problem is somewhere thinking harder doesn't reach.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
high1224 ±44-77.0$0.0600.7 min9,956#123
default (no effort flag)1200 ±41-76.9$0.0320.4 min2,501#126

The 95% Elo ranges overlap: Gemini 3.8 Flash (default) runs from 1157 to 1238, and Gemini 3.8 Flash (high thinking) from 1185 to 1273.

Which setting should you use?

Save the money and stay on the default. A gap that size is exactly the kind of thing that moves around between runs.

What more thinking doesn't fix is length. Both settings land far further from the target script length than Gemini 3.7 Flash (default) does, and that's the weak spot I'd watch here. If you need scripts that run the right length, I'd start with that one instead: it lands in the same Elo range and costs $0.028 a script.

Other Google ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Google model side by side.