DeepSeek · open weights · writing benchmark
DeepSeek V4.1 Flash (high) lands in the upper third of ToneBench at #45 of 181, with a writing Elo of 2066 and an overall score of 86.8 out of 100. Those numbers come from it writing its own version of the scripts behind 10 of my What's AI videos, five times each, with every draft scored blind against mine (how the scoring works).
Its Elo comes from comparing its average on each script with every other model's average on the same script, with gaps inside the run-to-run noise counted as draws. The 95% range is 2021 to 2109. That range overlaps 19 other configurations, ranked #31 to #50, so its exact spot inside that band is too close to call on this board. What the ranges can and can't tell you. One disclosure: the DeepSeek V4.1 Flash line also holds one of the four judge seats. The other three judges come from other model families, so no family scores itself alone (how the panel works).
Its biggest edge is Length: 87.6, 47th on the board and about 16 points above the board average. Even its weakest, Cues (79.9), sits above the board average, so there's no real weak spot to plan around.
The format changes the result. Its best score was on the career guide / hiring analysis (89.2) and its weakest on the engineering process walkthrough / presentation adaptation (84.7), about 4 points apart. If most of your scripts look like one of those, weigh that score over the average.
It costs about $0.009 per script. MiMo-V2.6-Flash scores higher and costs about $0.004 per script, so on value this isn't the pick. Still, plenty of pricier configs score lower than it does.
If you're on DeepSeek anyway, DeepSeek V4.1 Flash (max) is worth paying under a cent more per script.
Among open-weights models, it's 8th of 56. The best open-weights config is GLM-5.3 at #12, and everything above it is closed, so open weights still trail the frontier on this kind of writing.
DeepSeek V4.1 Flash is also on the board at other effort settings, and DeepSeek V4.1 Flash (max) writes better, at Elo 2165. Check the thinking-levels page before you lock this one in: it shows what each setting costs and how long it takes.
Read the shape: the further a corner reaches, the stronger the model is on that metric. The dashed line is the board average and the dotted line is the board's best score on each metric, so a corner outside the dashed line beats the average there. Hover a point for the exact numbers.
Nine writing metrics, each scored 0–100 and blended into the overall by the editorial weight shown next to it. Rank is out of all 181 current-ranked models. The thick bar is this model; the thin line above it is the field's best on that metric (darker) and the one below is the field's average (lighter), on the same scale. The small ± is how much the score moves from run to run, and the ? next to each name says what the metric rewards.
The same 10 real scripts every current-ranked model writes, each scored on its own as the mean of 5 runs. Different formats stress different skills: a model can nail a tight explainer and still stumble on a personal story, so look for the format closest to what you write.
A model whose drafts swing from run to run is harder to use than its average suggests. Every score on this page is the mean of 5 runs per script, and the ± is the run-to-run standard deviation. Overall, DeepSeek V4.1 Flash (high) moves by ± 1.8 against a board median of ± 2.5, so it's steadier than the typical model here.
| Metric | This model | Board median | Verdict |
|---|---|---|---|
| Tone ? | 87.0 ± 1.4 | ± 1.5 | typical |
| Craft ? | 88.7 ± 0.9 | ± 1.4 | typical |
| Substance ? | 86.1 ± 1.5 | ± 1.7 | typical |
| Flow ? | 86.2 ± 1.3 | ± 1.6 | typical |
| YouTube ? | 85.7 ± 2.1 | ± 2.5 | typical |
| Hook ? | 87.8 ± 1.6 | ± 1.9 | typical |
| Length ? | 87.6 ± 8.0 | ± 8.2 | typical |
| Slop ? | 89.8 ± 2.0 | ± 1.9 | typical |
| Cues ? | 79.9 ± 3.3 | ± 2.8 | typical |
Logged on every run. None of it touches the writing score, but it's often what decides whether a model fits your pipeline at all. Time is the average time to get one accepted script, retries included, and cost is priced per script at list price.
One judge can have taste of its own, which is why there are four: three LLM judges from three model families, plus Jev, a typed judge from TypeSafe, each scoring the same 50 stored drafts blind against the same rubric. The overall of 86.8 combines their scores with the parts computed in code (length, and half of the slop score). The highest and lowest judge are 3.6 points apart on its overall. That's tighter than the board median of 6.3, so the panel mostly agrees on this one. How the panel works.
deepseek/deepseek-v4.1-flash at reasoning_effort: highIt's #45 of 181 on ToneBench, with a writing Elo of 2066 and an overall score of 86.8/100. It wrote its own version of the scripts behind 10 of my What's AI videos 5 times each, and four judges scored every draft blind against mine: three LLM judges from three model families, plus Jev, a typed judge from TypeSafe. Length and half of the slop score are computed in code.
No. It's 3rd of 15 DeepSeek configs; DeepSeek V4.1 Flash (max) is about 99 Elo ahead.
It costs about $0.009 per script. MiMo-V2.6-Flash scores higher and costs less, so it isn't the value pick.
Against the board average, its biggest edge is Length (87.6, +15.6) and even its weakest, Cues (79.9, +2.3), is above average. The nine-metric breakdown on this page shows the board average and board best for each.
Through OpenRouter, using the exact model/route id deepseek/deepseek-v4.1-flash at reasoning_effort high, first run on 2026-09-29. Each script ran 5 times with provider-default sampling and no fine-tuning. The 'How we ran it' section and the methodology page have the rest.
Yes, it's an open-weights model. Among open-weights models it's 8th of 56 for writing.