Thinking levels · Google
Up to a point. Moving Gemini 3.1 Pro off low thinking is worth it: high beats it by 132 Elo and would win about 68% of their script comparisons, for 1.9 times the price and 1.4 minutes a script instead of 0.5. Past that, extra thinking stops paying, which is why the default is the setting to look at.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Reasoning tokens | Board rank |
|---|---|---|---|---|---|---|---|
| low | 1221 ±40 | - | 77.2 | $0.078 | 0.5 min | 2,286 | #124 |
| high | 1353 ±39 | +132 | 79.2 | $0.146 | 1.4 min | 7,422 | #112 |
| default (no effort flag) | 1384 ±44 | - | 79.4 | $0.143 | 1.3 min | 7,209 | #109 |
I'd leave Gemini 3.1 Pro on its default. It matches high thinking for about the same price, $0.143 against $0.146, and the two are close enough that the board can't separate them with confidence. So there's nothing to gain by forcing high.
Low thinking is the real trade-off. It's the cheap and fast option, but the 164 Elo you give up against the default is strong evidence of a weaker script, not noise. I'd only drop to low if you're generating at volume and every second counts.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Google model side by side.