Thinking levels · Google
Not on this board, and if anything it goes the wrong way. Turning Gemini 3.7 Flash up to high thinking makes every script cost more and take longer, and it still lands below the default setting. So the knob you'd normally reach for to get a better script is one I'd leave alone here.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Reasoning tokens | Board rank |
|---|---|---|---|---|---|---|---|
| high | 1169 ±39 | - | 76.7 | $0.043 | 0.6 min | 5,796 | #133 |
| default (no effort flag) | 1242 ±34 | - | 77.6 | $0.028 | 0.4 min | 1,951 | #121 |
The 95% Elo ranges overlap: Gemini 3.7 Flash (default) runs from 1206 to 1273, and Gemini 3.7 Flash (high thinking) from 1130 to 1208.
Stay on the default. High thinking costs 1.5 times as much, takes 0.6 minutes a script instead of 0.4, and lands 73 Elo lower, which puts it at #133 on the board against #121 for the default.
To be fair, the two ranges only just touch, so I can't call high thinking proven worse. But I can't find anything in the data that says it helps either, and you'd be paying extra to find out. The default already does the job for $0.028 a script, so why mess with it?
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Google model side by side.