Google won't give you the top script on this board. Its best, Gemini 3.1 Pro (default), sits at #109, a long way from the leaders. The reason to read on is the ladder below it: 17 cheap, fast configurations across 11 models, all closed except Gemma, with 1615 Elo between the top rung and the bottom one. And honestly, the names don't help you pick. A newer Flash isn't always a better Flash, so it's worth reading the table before you trust the version number.
Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.
| Board rank | Configuration | Elo | Overall | Tone | Craft | Substance | Hook | YouTube | Flow | Slop | Length | Cues | Cost / script | Time / script |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| #109 | Gemini 3.1 Pro (default) | 1384 ±44 | 79.4 | 77.8 | 79.2 | 80.0 | 82.8 | 79.2 | 77.9 | 85.1 | 78.7 | 77.1 | $0.143 | 1.3 min |
| #112 | Gemini 3.1 Pro (high thinking) | 1353 ±39 | 79.2 | 77.4 | 79.2 | 79.6 | 82.9 | 79.2 | 77.9 | 85.4 | 77.5 | 76.6 | $0.146 | 1.4 min |
| #119 | Gemini 2.5 Pro | 1264 ±40 | 77.8 | 76.0 | 80.3 | 75.6 | 80.3 | 76.6 | 76.4 | 81.0 | 83.1 | 73.4 | $0.060 | 0.6 min |
| #121 | Gemini 3.7 Flash (default) | 1242 ±34 | 77.6 | 72.6 | 77.8 | 77.2 | 82.0 | 76.6 | 76.2 | 79.9 | 82.5 | 85.8 | $0.028 | 0.4 min |
| #123 | Gemini 3.8 Flash (high thinking) | 1224 ±44 | 77.0 | 76.5 | 80.3 | 78.8 | 82.7 | 79.3 | 78.3 | 79.9 | 49.7 | 85.9 | $0.060 | 0.7 min |
| #124 | Gemini 3.1 Pro (low thinking) | 1221 ±40 | 77.2 | 75.9 | 78.0 | 77.3 | 81.5 | 76.9 | 76.2 | 84.8 | 70.0 | 78.4 | $0.078 | 0.5 min |
| #126 | Gemini 3.8 Flash (default) | 1200 ±41 | 76.9 | 75.7 | 79.7 | 78.4 | 83.1 | 78.9 | 78.3 | 79.3 | 52.9 | 86.2 | $0.032 | 0.4 min |
| #128 | Gemini 3.6 Flash (high thinking) | 1197 ±45 | 76.8 | 72.4 | 75.4 | 76.6 | 81.8 | 77.0 | 74.4 | 78.2 | 84.2 | 82.7 | $0.050 | 0.9 min |
| #129 | Gemini 3.6 Flash (default) | 1194 ±42 | 76.8 | 72.5 | 75.8 | 76.7 | 81.8 | 75.8 | 74.6 | 78.5 | 83.2 | 83.2 | $0.043 | 0.8 min |
| #133 | Gemini 3.7 Flash (high thinking) | 1169 ±39 | 76.7 | 73.2 | 77.1 | 77.9 | 82.2 | 77.7 | 75.5 | 78.5 | 69.6 | 85.9 | $0.043 | 0.6 min |
| #139 | Gemini 3.5 Flash (default) | 1087 ±42 | 75.2 | 72.7 | 74.5 | 74.3 | 80.2 | 75.1 | 73.1 | 79.0 | 77.2 | 80.1 | $0.142 | 1.4 min |
| #141 | Gemini 3.5 Flash (high thinking) | 1074 ±55 | 74.5 | 72.9 | 74.4 | 73.6 | 79.7 | 76.5 | 73.4 | 79.0 | 66.5 | 81.3 | $0.170 | 1.4 min |
| #147 | Gemini 2.5 Flash | 952 ±82 | 71.0 | 71.0 | 72.7 | 76.8 | 74.2 | 66.8 | 67.4 | 73.0 | 62.5 | 75.0 | $0.015 | 0.4 min |
| #148 | Gemini 3.5 Flash-Lite | 901 ±46 | 71.8 | 67.8 | 75.7 | 73.9 | 77.4 | 67.3 | 72.6 | 79.7 | 63.6 | 74.5 | $0.0085 | 0.1 min |
| #156 | Gemini 3.1 Flash-Lite | 769 ±46 | 69.7 | 70.1 | 75.3 | 68.8 | 76.3 | 63.2 | 70.2 | 79.5 | 56.3 | 68.2 | $0.0054 | 0.1 min |
| #161 | Gemini 2.5 Flash-Lite | 667 ±41 | 66.8 | 66.0 | 71.6 | 73.3 | 67.7 | 56.0 | 66.9 | 67.7 | 62.2 | 67.8 | $0.0021 | 0.2 min |
| #180 | Gemma 3 4B · open | -230 ±39 | 45.2 | 46.6 | 50.0 | 41.7 | 46.8 | 39.5 | 43.0 | 59.2 | 46.5 | 36.4 | $0.0008 | 0.7 min |
At the top, Gemini 3.1 Pro (default) edges out its own high-thinking setting, though by too little for me to read much into the order. Either way, you don't need the extra thinking to get Google's best script.
The one I'd run is Gemini 3.7 Flash (default). Gemini 3.1 Pro (default) costs 5.0 times as much for 142 more Elo, and at $0.028 a script you can generate a lot of drafts before it adds up. It isn't even the newest Flash: Gemini 3.8 Flash (default) came later and doesn't beat it here.
That price is also why I wouldn't go lower. Gemma 3 4B is the cheapest thing here at $0.0008, and it's last in the family at #180. The Flash-Lite models cost less than the value pick too, but they score far enough below it that I wouldn't trade the quality for the savings.
One to skip: Gemini 3.5 Flash (default) costs 5.0 times what Gemini 3.7 Flash (default) does and scores lower. I can't find a reason to pick it today.
Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.
Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.
← All comparisons