Towards AITowards AIToneBench

Qwen3 235B A22B

Alibaba (Qwen) · open weights · writing benchmark

#141 of 167 Elo 830 Overall 68.5 Open weights
Rank
#141 of 167
Writing Elo
830 ±60
Overall
68.5 / 100
Cost / script
$0.003 per script
Family
Alibaba (Qwen) 6th of 10
Type
Open 33rd of 53
Consistency
± 5.0 swingier than most
Avg tokens
14.8k in+out
Time / script
0.8 min avg, retries included
Family check:Qwen3.8 Flash is Alibaba (Qwen)'s best writer here: +14.5 overall against this config for $0.010 more per script.

The short version

Qwen3 235B A22B lands in the bottom third of ToneBench at #141 of 167, with a writing Elo of 830 and an overall score of 68.5 out of 100. Those numbers come from it writing its own version of the scripts behind 10 of my What's AI videos, five times each, with every draft scored blind against mine (how the scoring works).

For scale, my own scripts score 90.9 through the same judges and rubric. That's the real bar here, not a literal 100. It sits about 22 points under it.

Its Elo comes from comparing its average on each script with every other model's average on the same script, with gaps inside the run-to-run noise counted as draws. The 95% range is 771 to 892. That range overlaps 12 other configurations, ranked #136 to #148, so its exact spot inside that band is too close to call on this board. What the ranges can and can't tell you.

It's under the board average on every metric. The least-bad is Length, still 2.7 points under. Where it slips is Cues (53.0, about 24 points under the board average): you'll be adding most of the visual cues yourself.

The format changes the result. Its best score was on the product release announcement / personal observation (74.3) and its weakest on the short explainer (60.3), about 14 points apart. If most of your scripts look like one of those, weigh that score over the average.

It's also swingy: its overall moves ± 5.0 between runs of the same script, against ± 2.7 for the typical model, so the same prompt can give you a usable draft one time and one you'd rewrite the next. If you use it, generate a few drafts and keep the best.

It costs about $0.003 per script. DeepSeek V4 Flash (default) scores far higher and costs about $0.002 per script, so on value this isn't the pick. Still, a few pricier configs score lower than it does.

If you're on Alibaba (Qwen) anyway, Qwen3.8 Flash is worth paying about 1 cents more per script.

Among open-weights models, it's 33rd of 53. The best open-weights config is GLM-5.3 at #15, and everything above it is closed, so open weights still trail the frontier on this kind of writing.

Skill profile

Read the shape: the further a corner reaches, the stronger the model is on that metric. The dashed line is the board average and the dotted line is the board's best score on each metric, so a corner outside the dashed line beats the average there. Hover a point for the exact numbers.

50Tone?Craft?Substance?Flow?YouTube?Hook?Length?Slop?Cues?
Qwen3 235B A22BBoard averageBoard best per metricMax possible (100)

Metric by metric

Nine writing metrics, each scored 0–100 and blended into the overall by the editorial weight shown next to it. Rank is out of all 167 current-ranked models. The thick bar is this model; the thin line above it is the field's best on that metric (darker) and the one below is the field's average (lighter), on the same scale. The small ± is how much the score moves from run to run, and the ? next to each name says what the metric rewards.

Tone ?19% weight · -6.1 vs avg
best · Claude Opus 5.5 (max effort) · 91.072.5± 2.8 · 134thboard avg · 78.6
Craft ?13% weight · -7.8 vs avg
best · Claude Opus 5.5 (max effort) · 91.572.8± 3.4 · 142ndboard avg · 80.5
Substance ?15% weight · -13.0 vs avg
best · Claude Opus 5.5 (xhigh) · 90.766.3± 5.0 · 148thboard avg · 79.4
Flow ?14% weight · -9.3 vs avg
best · Claude Opus 5.5 (max effort) · 90.967.7± 5.5 · 139thboard avg · 77.0
YouTube ?12% weight · -15.3 vs avg
best · Claude Opus 5.5 (max effort) · 91.761.4± 5.1 · 150thboard avg · 76.7
Hook ?10% weight · -4.2 vs avg
best · Claude Opus 5.5 (max effort) · 91.277.4± 4.5 · 140thboard avg · 81.6
Length ?8% weight · -2.7 vs avg
best · Claude Opus 5.5 (max effort) · 97.668.1± 17.4 · 103rdboard avg · 70.8
Slop ?5% weight · -21.3 vs avg
best · Claude Opus 5.5 (max effort) · 94.962.9± 2.0 · 164thboard avg · 84.2
Cues ?4% weight · -24.0 vs avg
best · GPT-5.6 Sol (ultra) · 90.153.0± 7.2 · 157thboard avg · 77.0

Script by script

The same 10 real scripts every current-ranked model writes, each scored on its own as the mean of 5 runs. Different formats stress different skills: a model can nail a tight explainer and still stumble on a personal story, so look for the format closest to what you write.

Script 1
opinion / warning explainer
68.9
out of 100
Script 2
news-analysis / skeptical explainer
68.3
out of 100
Script 3
personal roadmap / opinion
66.8
out of 100
Script 4
short explainer
60.3
out of 100
Script 5
news-analysis / opinion explainer
67.8
out of 100
Script 6
founder announcement / personal origin story
72.0
out of 100
Script 7
personal technical walkthrough / agentic workflow case study
67.7
out of 100
Script 8
product release announcement / personal observation
74.3
out of 100
Script 9
career guide / hiring analysis
71.3
out of 100
Script 10
engineering process walkthrough / presentation adaptation
67.2
out of 100

Consistency: how much it moves between runs

A model whose drafts swing from run to run is harder to use than its average suggests. Every score on this page is the mean of 5 runs per script, and the ± is the run-to-run standard deviation.

Its most volatile metric is Flow: ± 5.5 between runs, when the typical model moves ± 1.8, so two drafts of the same script can score quite differently there.

MetricThis modelBoard medianVerdict
Tone ?72.5 ± 2.8± 1.7swingier than most
Craft ?72.8 ± 3.4± 1.5swingier than most
Substance ?66.3 ± 5.0± 2.0swingier than most
Flow ?67.7 ± 5.5± 1.8swingier than most
YouTube ?61.4 ± 5.1± 2.9swingier than most
Hook ?77.4 ± 4.5± 2.0swingier than most
Length ?68.1 ± 17.4± 8.6swingier than most
Slop ?62.9 ± 2.0± 2.1typical
Cues ?53.0 ± 7.2± 3.2swingier than most

Cost, time and tokens

Logged on every run. None of it touches the writing score, but it's often what decides whether a model fits your pipeline at all. Time is the average time to get one accepted script, retries included, and cost is priced per script at list price.

Time per script
0.8 minfaster than the board median of 1.2 min
Prompt tokens in
11.6kstyle guide + brief + research packet
Tokens out
3.1kscript + any reasoning tokens
Cost per script
$0.003 ± 0.001measured from actual billed tokens
List price used
$0.09 / $0.55 per M tokinput / output

What each judge scored it

One judge can have taste of its own, which is why there are three, from three model families, each scoring the same 50 stored drafts blind with the identical rubric. The overall of 68.5 combines their scores with the parts computed in code (length, and half of the slop score). The highest and lowest judge are 10.0 points apart on its overall. That's wider than typical (board median 6.3): DeepSeek V4.1 Flash rates it noticeably higher than Claude Opus 5, so the consensus sits between two real opinions. How the panel works.

Claude Opus 5
Anthropic
63.4
GPT-5.6 Sol (medium)
OpenAI
68.6
DeepSeek V4.1 Flash
DeepSeek
73.4

How we ran Qwen3 235B A22B

Frequently asked questions

How good is Qwen3 235B A22B at writing?

It's #141 of 167 on ToneBench, with a writing Elo of 830 and an overall score of 68.5/100. It wrote its own version of the scripts behind 10 of my What's AI videos 5 times each, and three judges from three model families scored every draft blind against mine. Length and half of the slop score are computed in code.

Is Qwen3 235B A22B the best Alibaba (Qwen) model for writing?

No. It's 6th of 10 Alibaba (Qwen) configs; Qwen3.8 Flash is about 869 Elo ahead.

Is Qwen3 235B A22B good value for the money?

It costs about $0.003 per script. DeepSeek V4 Flash (default) scores higher and costs less, so it isn't the value pick.

What are Qwen3 235B A22B's strengths and weaknesses?

It's under the board average on every metric; the least-bad is Length (68.1, -2.7) and it's furthest behind on Cues (53.0, -24.0). The nine-metric breakdown on this page shows the board average and board best for each.

How was Qwen3 235B A22B evaluated?

Through OpenRouter, using the exact model/route id qwen/qwen3-235b-a22b-2507, first run on 2026-09-03. Each script ran 5 times with provider-default sampling and no fine-tuning. The 'How we ran it' section and the methodology page have the rest.

Is Qwen3 235B A22B open source?

Yes, it's an open-weights model. Among open-weights models it's 33rd of 53 for writing.

← Back to the full ToneBench leaderboard