Towards AITowards AIToneBench

Claude Sonnet 5.5 (max effort)

Anthropic · proprietary · writing benchmark

#3 of 181 Elo 2512 Overall 90.1 Proprietary
Rank
#3 of 181
Writing Elo
2512 ±41
Overall
90.1 / 100
Cost / script
$1.32 per script
Family
Anthropic 3rd of 35
Type
Closed 3rd of 125
Consistency
± 1.3 steadier than most
Avg tokens
141.4k in+out
Time / script
12.9 min avg, retries included
Family check:Claude Opus 5.5 (max effort) is Anthropic's best writer here: +0.8 overall against this config for $1.422 more per script.

The short version

Claude Sonnet 5.5 (max effort) is one of the handful of models right at the top of ToneBench: #3 of 181, with a writing Elo of 2512 and an overall score of 90.1 out of 100. Those numbers come from it writing its own version of the scripts behind 10 of my What's AI videos, five times each, with every draft scored blind against mine (how the scoring works).

Its Elo comes from comparing its average on each script with every other model's average on the same script, with gaps inside the run-to-run noise counted as draws. The 95% range is 2470 to 2553. That range overlaps 2 other configurations, ranked #2 to #4, so its exact spot inside that band is too close to call on this board. What the ranges can and can't tell you.

Its biggest edge is Length: 98.2, 1st on the board and about 26 points above the board average. Even its weakest, Hook (89.6), sits above the board average, so there's no real weak spot to plan around.

The format changes the result. Its best score was on the short explainer (91.8) and its weakest on the founder announcement / personal origin story, the product release announcement / personal observation and the engineering process walkthrough / presentation adaptation (all 88.8), about 3 points apart. If most of your scripts look like one of those, weigh that score over the average.

It's also unusually consistent: its overall moves only ± 1.3 between runs of the same script, against ± 2.5 for the typical model. What you get on the first try is close to what you get every time.

It costs about $1.32 per script. Claude Opus 5.5 (xhigh) scores higher and costs about $0.69 per script, so on value this isn't the pick.

If you're on Anthropic anyway, Claude Opus 5.5 (max effort) is worth paying about $1.42 more per script.

Among proprietary models, it's 3rd of 125. It also sits above every open-weights model; the best of them, GLM-5.3, is #12.

Claude Sonnet 5.5 is also on the board at other effort settings, and none of them beat this one; the closest is Claude Sonnet 5.5 (xhigh), at Elo 2303. The thinking-levels page shows what each setting costs and how long it takes, so you can see how much quality a cheaper or faster one gives up.

Skill profile

Read the shape: the further a corner reaches, the stronger the model is on that metric. The dashed line is the board average and the dotted line is the board's best score on each metric, so a corner outside the dashed line beats the average there. Hover a point for the exact numbers.

50Tone?Craft?Substance?Flow?YouTube?Hook?Length?Slop?Cues?
Claude Sonnet 5.5 (max effort)Board averageBoard best per metricMax possible (100)

Metric by metric

Nine writing metrics, each scored 0–100 and blended into the overall by the editorial weight shown next to it. Rank is out of all 181 current-ranked models. The thick bar is this model; the thin line above it is the field's best on that metric (darker) and the one below is the field's average (lighter), on the same scale. The small ± is how much the score moves from run to run, and the ? next to each name says what the metric rewards.

Tone ?19% weight · +10.1 vs avg
best · Claude Opus 5.5 (max effort) · 90.489.3± 0.9 · 12thboard avg · 79.1
Craft ?13% weight · +8.7 vs avg
best · Claude Opus 5.5 (max effort) · 90.989.8± 0.7 · 10thboard avg · 81.1
Substance ?15% weight · +9.1 vs avg
best · Claude Opus 5.5 (max effort) · 89.489.1± 0.9 · 3rdboard avg · 80.0
Flow ?14% weight · +10.7 vs avg
best · Claude Opus 5.5 (max effort) · 89.588.4± 1.0 · 8thboard avg · 77.7
YouTube ?12% weight · +11.8 vs avg
best · Claude Opus 5.5 (max effort) · 90.189.3± 1.1 · 5thboard avg · 77.5
Hook ?10% weight · +7.6 vs avg
best · Claude Opus 5.5 (max effort) · 90.489.6± 1.6 · 4thboard avg · 82.0
Length ?8% weight · +26.2 vs avg
best on the board · this model · 98.298.2± 0.9 · 1stboard avg · 72.0
Slop ?5% weight · +8.8 vs avg
best · GPT-6 Astra (max) · 94.893.6± 0.9 · 11thboard avg · 84.8
Cues ?4% weight · +9.0 vs avg
best · GPT-5.6 Sol (ultra) · 89.586.6± 2.1 · 31stboard avg · 77.6

Script by script

The same 10 real scripts every current-ranked model writes, each scored on its own as the mean of 5 runs. Different formats stress different skills: a model can nail a tight explainer and still stumble on a personal story, so look for the format closest to what you write.

Script 1
opinion / warning explainer
90.4
out of 100
Script 2
news-analysis / skeptical explainer
90.0
out of 100
Script 3
personal roadmap / opinion
90.5
out of 100
Script 4
short explainer
91.8
out of 100
Script 5
news-analysis / opinion explainer
89.5
out of 100
Script 6
founder announcement / personal origin story
88.8
out of 100
Script 7
personal technical walkthrough / agentic workflow case study
90.5
out of 100
Script 8
product release announcement / personal observation
88.8
out of 100
Script 9
career guide / hiring analysis
91.6
out of 100
Script 10
engineering process walkthrough / presentation adaptation
88.8
out of 100

Consistency: how much it moves between runs

A model whose drafts swing from run to run is harder to use than its average suggests. Every score on this page is the mean of 5 runs per script, and the ± is the run-to-run standard deviation.

Its steadiest is Length: ± 0.9 run to run, when the typical model swings ± 8.2.

MetricThis modelBoard medianVerdict
Tone ?89.3 ± 0.9± 1.5steadier than most
Craft ?89.8 ± 0.7± 1.4steadier than most
Substance ?89.1 ± 0.9± 1.7steadier than most
Flow ?88.4 ± 1.0± 1.6typical
YouTube ?89.3 ± 1.1± 2.5steadier than most
Hook ?89.6 ± 1.6± 1.9typical
Length ?98.2 ± 0.9± 8.2steadier than most
Slop ?93.6 ± 0.9± 1.9steadier than most
Cues ?86.6 ± 2.1± 2.8typical

Cost, time and tokens

Logged on every run. None of it touches the writing score, but it's often what decides whether a model fits your pipeline at all. Time is the average time to get one accepted script, retries included, and cost is priced per script at list price.

Time per script
12.9 minslower than the board median of 1.2 min
Prompt tokens in
11.3kstyle guide + brief + research packet
Tokens out
130.1kwell above the board median (a lot of thinking)
Cost per script
$1.324 ± 0.350output measured incl. thinking; input priced as the prompt billed once
List price used
$2 / $10 per M tokinput / output

What each judge scored it

One judge can have taste of its own, which is why there are four: three LLM judges from three model families, plus Jev, a typed judge from TypeSafe, each scoring the same 50 stored drafts blind against the same rubric. The overall of 90.1 combines their scores with the parts computed in code (length, and half of the slop score). The highest and lowest judge are 6.2 points apart on its overall. That's a typical amount of judge disagreement for this board. How the panel works.

Claude Opus 5
Anthropic
89.6
GPT-5.6 Sol (medium)
OpenAI
93.9
DeepSeek V4.1 Flash
DeepSeek
89.1
Jev 1.13
TypeSafe
87.7

How we ran Claude Sonnet 5.5 (max effort)

Frequently asked questions

How good is Claude Sonnet 5.5 (max effort) at writing?

It's #3 of 181 on ToneBench, with a writing Elo of 2512 and an overall score of 90.1/100. It wrote its own version of the scripts behind 10 of my What's AI videos 5 times each, and four judges scored every draft blind against mine: three LLM judges from three model families, plus Jev, a typed judge from TypeSafe. Length and half of the slop score are computed in code.

Is Claude Sonnet 5.5 (max effort) the best Anthropic model for writing?

No. It's 3rd of 35 Anthropic configs; Claude Opus 5.5 (max effort) is about 171 Elo ahead.

Is Claude Sonnet 5.5 (max effort) good value for the money?

It costs about $1.32 per script. Claude Opus 5.5 (xhigh) scores higher and costs less, so it isn't the value pick.

What are Claude Sonnet 5.5 (max effort)'s strengths and weaknesses?

Against the board average, its biggest edge is Length (98.2, +26.2) and even its weakest, Hook (89.6, +7.6), is above average. The nine-metric breakdown on this page shows the board average and board best for each.

How was Claude Sonnet 5.5 (max effort) evaluated?

Through Claude Code (Anthropic subscription), using the exact model/route id claude-sonnet-5-5 at reasoning_effort max, first run on 2026-09-28. Each script ran 5 times with provider-default sampling and no fine-tuning. The 'How we ran it' section and the methodology page have the rest.

Is Claude Sonnet 5.5 (max effort) open source?

No, it's a proprietary (closed-weights) model. Among proprietary models it's 3rd of 125 for writing.

← Back to the full ToneBench leaderboard