Towards AITowards AIToneBench

Thinking levels · Google

Does Gemini 3.6 Flash write better when it thinks harder?

Hardly at all. On Gemini 3.6 Flash, high thinking and the default end up 3 Elo apart, which is nothing on this scale. You still pay for the extra thinking, $0.050 a script against $0.043, so the only difference you'd actually notice is on the bill.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
high1197 ±45-76.8$0.0500.9 min8,009#128
default (no effort flag)1194 ±42-76.8$0.0430.8 min5,984#129

The 95% Elo ranges overlap: Gemini 3.6 Flash (default) runs from 1151 to 1234, and Gemini 3.6 Flash (high thinking) from 1148 to 1238.

Which setting should you use?

Default, without much hesitation. High thinking costs 1.2 times as much and takes a bit longer. It's a small premium, but I can't see what it buys you.

One thing that holds at either setting: Gemini 3.6 Flash hits the target length well, which the newer Gemini 3.8 Flash doesn't. That said, I'd still look at Gemini 3.7 Flash (default) first. It's cheaper at $0.028 and scores at least as well.

Other Google ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Google model side by side.