Towards AITowards AIToneBench

Gemini 3.7 Flash (high thinking)

Google · proprietary · writing benchmark

#133 of 181 Elo 1169 Overall 76.7 Proprietary
Rank
#133 of 181
Writing Elo
1169 ±39
Overall
76.7 / 100
Cost / script
$0.043 per script
Family
Google 10th of 17
Type
Closed 103rd of 125
Consistency
± 2.5 typical spread
Avg tokens
21.1k in+out
Time / script
0.6 min avg, retries included
Family check:Gemini 3.1 Pro (default) is Google's best writer here: +2.8 overall against this config for $0.100 more per script.

The short version

Gemini 3.7 Flash (high thinking) lands in the bottom third of ToneBench at #133 of 181, with a writing Elo of 1169 and an overall score of 76.7 out of 100. Those numbers come from it writing its own version of the scripts behind 10 of my What's AI videos, five times each, with every draft scored blind against mine (how the scoring works).

Its Elo comes from comparing its average on each script with every other model's average on the same script, with gaps inside the run-to-run noise counted as draws. The 95% range is 1130 to 1208. That range overlaps 18 other configurations, ranked #121 to #139, so its exact spot inside that band is too close to call on this board. What the ranges can and can't tell you.

Its biggest edge is Cues: 85.9, 48th on the board and 8.2 points above the board average. Where it slips is Slop (higher means cleaner) (78.5, 6.2 points under the board average): expect filler phrases and AI tells to clean out.

The format changes the result. Its best score was on the news-analysis / skeptical explainer (79.6) and its weakest on the engineering process walkthrough / presentation adaptation (72.1), about 8 points apart. If most of your scripts look like one of those, weigh that score over the average.

It costs about $0.043 per script. DeepSeek V4 Flash (native reasoner alias) scores higher and costs about $0.002 per script, so on value this isn't the pick. Still, a few pricier configs score lower than it does.

If you're on Google anyway, Gemini 3.1 Pro (default) is worth paying about 10 cents more per script.

Among proprietary models, it's 103rd of 125. The best open-weights config, GLM-5.3 at #12, outranks it, so it's worth a look if you'd rather run open weights.

Gemini 3.7 Flash is also on the board at other effort settings, and Gemini 3.7 Flash (default) writes better, at Elo 1242. Check the thinking-levels page before you lock this one in: it shows what each setting costs and how long it takes.

Skill profile

Read the shape: the further a corner reaches, the stronger the model is on that metric. The dashed line is the board average and the dotted line is the board's best score on each metric, so a corner outside the dashed line beats the average there. Hover a point for the exact numbers.

50Tone?Craft?Substance?Flow?YouTube?Hook?Length?Slop?Cues?
Gemini 3.7 Flash (high thinking)Board averageBoard best per metricMax possible (100)

Metric by metric

Nine writing metrics, each scored 0–100 and blended into the overall by the editorial weight shown next to it. Rank is out of all 181 current-ranked models. The thick bar is this model; the thin line above it is the field's best on that metric (darker) and the one below is the field's average (lighter), on the same scale. The small ± is how much the score moves from run to run, and the ? next to each name says what the metric rewards.

Tone ?19% weight · -5.9 vs avg
best · Claude Opus 5.5 (max effort) · 90.473.2± 2.3 · 141stboard avg · 79.1
Craft ?13% weight · -4.0 vs avg
best · Claude Opus 5.5 (max effort) · 90.977.1± 2.0 · 139thboard avg · 81.1
Substance ?15% weight · -2.1 vs avg
best · Claude Opus 5.5 (max effort) · 89.477.9± 1.6 · 126thboard avg · 80.0
Flow ?14% weight · -2.2 vs avg
best · Claude Opus 5.5 (max effort) · 89.575.5± 1.5 · 128thboard avg · 77.7
YouTube ?12% weight · +0.2 vs avg
best · Claude Opus 5.5 (max effort) · 90.177.7± 2.1 · 114thboard avg · 77.5
Hook ?10% weight · +0.2 vs avg
best · Claude Opus 5.5 (max effort) · 90.482.2± 1.9 · 117thboard avg · 82.0
Length ?8% weight · -2.4 vs avg
best · Claude Sonnet 5.5 (max effort) · 98.269.6± 15.2 · 115thboard avg · 72.0
Slop ?5% weight · -6.2 vs avg
best · GPT-6 Astra (max) · 94.878.5± 2.3 · 145thboard avg · 84.8
Cues ?4% weight · +8.2 vs avg
best · GPT-5.6 Sol (ultra) · 89.585.9± 2.2 · 48thboard avg · 77.6

Script by script

The same 10 real scripts every current-ranked model writes, each scored on its own as the mean of 5 runs. Different formats stress different skills: a model can nail a tight explainer and still stumble on a personal story, so look for the format closest to what you write.

Script 1
opinion / warning explainer
75.7
out of 100
Script 2
news-analysis / skeptical explainer
79.6
out of 100
Script 3
personal roadmap / opinion
78.2
out of 100
Script 4
short explainer
75.9
out of 100
Script 5
news-analysis / opinion explainer
76.5
out of 100
Script 6
founder announcement / personal origin story
77.6
out of 100
Script 7
personal technical walkthrough / agentic workflow case study
75.5
out of 100
Script 8
product release announcement / personal observation
78.4
out of 100
Script 9
career guide / hiring analysis
76.8
out of 100
Script 10
engineering process walkthrough / presentation adaptation
72.1
out of 100

Consistency: how much it moves between runs

A model whose drafts swing from run to run is harder to use than its average suggests. Every score on this page is the mean of 5 runs per script, and the ± is the run-to-run standard deviation. Overall, Gemini 3.7 Flash (high thinking) moves by ± 2.5 against a board median of ± 2.5, so it's about as repeatable as the typical model on the board.

Its most volatile metric is Length: ± 15.2 between runs, when the typical model moves ± 8.2, so two drafts of the same script can score quite differently there.

MetricThis modelBoard medianVerdict
Tone ?73.2 ± 2.3± 1.5swingier than most
Craft ?77.1 ± 2.0± 1.4typical
Substance ?77.9 ± 1.6± 1.7typical
Flow ?75.5 ± 1.5± 1.6typical
YouTube ?77.7 ± 2.1± 2.5typical
Hook ?82.2 ± 1.9± 1.9typical
Length ?69.6 ± 15.2± 8.2swingier than most
Slop ?78.5 ± 2.3± 1.9typical
Cues ?85.9 ± 2.2± 2.8typical

Cost, time and tokens

Logged on every run. None of it touches the writing score, but it's often what decides whether a model fits your pipeline at all. Time is the average time to get one accepted script, retries included, and cost is priced per script at list price.

Time per script
0.6 minfaster than the board median of 1.2 min
Prompt tokens in
12.0kstyle guide + brief + research packet
Tokens out
9.1kwell above the board median (a lot of thinking)
Cost per script
$0.043 ± 0.010measured from actual billed tokens
List price used
$0.75 / $3.75 per M tokinput / output

What each judge scored it

One judge can have taste of its own, which is why there are four: three LLM judges from three model families, plus Jev, a typed judge from TypeSafe, each scoring the same 50 stored drafts blind against the same rubric. The overall of 76.7 combines their scores with the parts computed in code (length, and half of the slop score). The highest and lowest judge are 5.7 points apart on its overall. That's a typical amount of judge disagreement for this board. How the panel works.

Claude Opus 5
Anthropic
74.0
GPT-5.6 Sol (medium)
OpenAI
79.7
DeepSeek V4.1 Flash
DeepSeek
77.5
Jev 1.13
TypeSafe
75.4

How we ran Gemini 3.7 Flash (high thinking)

Frequently asked questions

How good is Gemini 3.7 Flash (high thinking) at writing?

It's #133 of 181 on ToneBench, with a writing Elo of 1169 and an overall score of 76.7/100. It wrote its own version of the scripts behind 10 of my What's AI videos 5 times each, and four judges scored every draft blind against mine: three LLM judges from three model families, plus Jev, a typed judge from TypeSafe. Length and half of the slop score are computed in code.

Is Gemini 3.7 Flash (high thinking) the best Google model for writing?

No. It's 10th of 17 Google configs; Gemini 3.1 Pro (default) is about 215 Elo ahead.

Is Gemini 3.7 Flash (high thinking) good value for the money?

It costs about $0.043 per script. DeepSeek V4 Flash (native reasoner alias) scores higher and costs less, so it isn't the value pick.

What are Gemini 3.7 Flash (high thinking)'s strengths and weaknesses?

Against the board average, its biggest edge is Cues (85.9, +8.2) and it's furthest behind on Slop (higher means cleaner) (78.5, -6.2). The nine-metric breakdown on this page shows the board average and board best for each.

How was Gemini 3.7 Flash (high thinking) evaluated?

Through Google Gemini API, using the exact model/route id gemini-3.7-flash at reasoning_effort high, first run on 2026-09-03. Each script ran 5 times with provider-default sampling and no fine-tuning. The 'How we ran it' section and the methodology page have the rest.

Is Gemini 3.7 Flash (high thinking) open source?

No, it's a proprietary (closed-weights) model. Among proprietary models it's 103rd of 125 for writing.

← Back to the full ToneBench leaderboard