Cohere · open weights · writing benchmark
Cohere Command R+ lands in the bottom third of ToneBench at #163 of 167, with a writing Elo of -33 and an overall score of 41.1 out of 100. Those numbers come from it writing its own version of the scripts behind 10 of my What's AI videos, five times each, with every draft scored blind against mine (how the scoring works).
For scale, my own scripts score 90.9 through the same judges and rubric. That's the real bar here, not a literal 100. It sits about 50 points under it.
Its Elo comes from comparing its average on each script with every other model's average on the same script, with gaps inside the run-to-run noise counted as draws. The 95% range is -105 to 47. That range overlaps 2 other configurations, ranked #162 to #165, so its exact spot inside that band is too close to call on this board. What the ranges can and can't tell you.
It's under the board average on every metric. The least-bad is Slop (higher means cleaner), still about 14 points under. It's furthest behind on Length, about 54 points under the board average.
The format changes the result. Its best score was on the product release announcement / personal observation (59.8) and its weakest on the founder announcement / personal origin story (14.1), about 46 points apart. If most of your scripts look like one of those, weigh that score over the average.
It's also swingy: its overall moves ± 18.5 between runs of the same script, against ± 2.7 for the typical model, so the same prompt can give you anything from rough to unusable.
It costs about $0.049 per script. Cohere Command A+ (05-2026) scores far higher and is free on the route we used, so on value this isn't the pick. For full scripts, I wouldn't use it.
If you're on Cohere anyway, Cohere Command A+ (05-2026) writes better for no more per script, so use that one.
Among open-weights models, it's 49th of 53. The best open-weights config is GLM-5.3 at #15, and everything above it is closed, so open weights still trail the frontier on this kind of writing.
Read the shape: the further a corner reaches, the stronger the model is on that metric. The dashed line is the board average and the dotted line is the board's best score on each metric, so a corner outside the dashed line beats the average there. Hover a point for the exact numbers.
Nine writing metrics, each scored 0–100 and blended into the overall by the editorial weight shown next to it. Rank is out of all 167 current-ranked models. The thick bar is this model; the thin line above it is the field's best on that metric (darker) and the one below is the field's average (lighter), on the same scale. The small ± is how much the score moves from run to run, and the ? next to each name says what the metric rewards.
The same 10 real scripts every current-ranked model writes, each scored on its own as the mean of 5 runs. Different formats stress different skills: a model can nail a tight explainer and still stumble on a personal story, so look for the format closest to what you write.
A model whose drafts swing from run to run is harder to use than its average suggests. Every score on this page is the mean of 5 runs per script, and the ± is the run-to-run standard deviation.
Its most volatile metric is Craft: ± 9.5 between runs, when the typical model moves ± 1.5, so two drafts of the same script can score quite differently there.
| Metric | This model | Board median | Verdict |
|---|---|---|---|
| Tone ? | 40.5 ± 8.8 | ± 1.7 | swingier than most |
| Craft ? | 46.8 ± 9.5 | ± 1.5 | swingier than most |
| Substance ? | 45.6 ± 9.4 | ± 2.0 | swingier than most |
| Flow ? | 39.0 ± 9.6 | ± 1.8 | swingier than most |
| YouTube ? | 33.3 ± 8.4 | ± 2.9 | swingier than most |
| Hook ? | 44.3 ± 11.5 | ± 2.0 | swingier than most |
| Length ? | 17.0 ± 10.1 | ± 8.6 | typical |
| Slop ? | 70.0 ± 6.7 | ± 2.1 | swingier than most |
| Cues ? | 42.6 ± 11.0 | ± 3.2 | swingier than most |
Logged on every run. None of it touches the writing score, but it's often what decides whether a model fits your pipeline at all. Time is the average time to get one accepted script, retries included, and cost is priced per script at list price.
One judge can have taste of its own, which is why there are three, from three model families, each scoring the same 50 stored drafts blind with the identical rubric. The overall of 41.1 combines their scores with the parts computed in code (length, and half of the slop score). The highest and lowest judge are 10.7 points apart on its overall. That's wider than typical (board median 6.3): GPT-5.6 Sol (medium) rates it noticeably higher than Claude Opus 5, so the consensus sits between two real opinions. How the panel works.
command-r-plus-08-2024 (provider default reasoning)It's #163 of 167 on ToneBench, with a writing Elo of -33 and an overall score of 41.1/100. It wrote its own version of the scripts behind 10 of my What's AI videos 5 times each, and three judges from three model families scored every draft blind against mine. Length and half of the slop score are computed in code.
No. It's 5th of 6 Cohere configs; Cohere Command A+ (05-2026) is about 544 Elo ahead.
It costs about $0.049 per script. Cohere Command A+ (05-2026) scores higher and costs less, so it isn't the value pick.
It's under the board average on every metric; the least-bad is Slop (higher means cleaner) (70.0, -14.2) and it's furthest behind on Length (17.0, -53.8). The nine-metric breakdown on this page shows the board average and board best for each.
Through Cohere API, using the exact model/route id command-r-plus-08-2024, first run on 2026-08-09. Each script ran 5 times with provider-default sampling and no fine-tuning. The 'How we ran it' section and the methodology page have the rest.
Yes, it's an open-weights model. Among open-weights models it's 49th of 53 for writing.