I'll be upfront: you won't find the board's strong script writers here. Mistral Medium 3.5 (reasoning) is the best Mistral at #151, and that's a bit of a shame, because Mistral matters to a lot of builders and shipping Mistral Large 3 as open weights is great. The 4 configurations are also packed close together, only 120 Elo from top to bottom, so the choice here comes down to price more than quality.
Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.
| Board rank | Configuration | Elo | Overall | Tone | Craft | Substance | Hook | YouTube | Flow | Slop | Length | Cues | Cost / script | Time / script |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| #151 | Mistral Medium 3.5 (reasoning) | 876 ±73 | 70.1 | 72.1 | 70.8 | 72.8 | 80.6 | 71.0 | 64.1 | 73.5 | 52.9 | 70.6 | $0.153 | 1.6 min |
| #152 | Mistral Medium 3.1 | 865 ±46 | 71.2 | 72.0 | 76.6 | 68.9 | 81.9 | 65.0 | 71.3 | 67.6 | 62.8 | 71.5 | $0.0094 | 0.5 min |
| #157 | Mistral Large 3 · open | 766 ±50 | 69.5 | 69.5 | 74.2 | 72.0 | 79.0 | 64.0 | 66.6 | 68.2 | 61.5 | 65.2 | $0.0090 | 1.0 min |
| #158 | Mistral Medium 3.5 | 756 ±45 | 69.5 | 67.6 | 74.5 | 73.2 | 78.8 | 65.9 | 67.8 | 77.9 | 46.3 | 76.9 | $0.031 | 0.3 min |
If you're set on Mistral, I'd pick Mistral Medium 3.1. Mistral Medium 3.5 (reasoning) tops the family, but only by 11 Elo with overlapping ranges, and it costs 16.3 times as much. Turn reasoning off and Mistral Medium 3.5 drops to #158, behind Mistral Medium 3.1 and still pricier, so that one is hard to justify. Mistral Large 3 is the cheapest at $0.0090, which makes it the value pick on paper, but Mistral Medium 3.1 scores higher for $0.0094, basically the same money. Length is the weak spot for all four, so plan on fixing it by hand. And if you're not tied to Mistral, I'd look higher up the board.
Only Mistral Medium 3.5 ran at more than one reasoning-effort setting here. Its page shows what each step up buys.
Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.
← All comparisons