Thinking levels · Mistral
Probably, but it's an expensive probably. Reasoning mode adds 120 Elo to Mistral Medium 3.5, which moves it from #158 to #151. The two ranges still overlap, though, so I wouldn't call it settled. It also costs 4.9 times as much per script and takes 1.6 minutes instead of 0.3.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Reasoning tokens | Board rank |
|---|---|---|---|---|---|---|---|
| high | 876 ±73 | - | 70.1 | $0.153 | 1.6 min | 15,228 | #151 |
| default (no effort flag) | 756 ±45 | - | 69.5 | $0.031 | 0.3 min | - | #158 |
The 95% Elo ranges overlap: Mistral Medium 3.5 runs from 709 to 800, and Mistral Medium 3.5 (reasoning) from 796 to 942.
I'd keep reasoning off. It probably helps, and by more than a little, but the board can't confirm the gain, and it multiplies the price for a draft that still sits low on the board. And the older Mistral Medium 3.1 gets about the same Elo as the reasoning mode for $0.0094 per script, no reasoning needed. In this family, reasoning mode is the priciest way to get that score.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Mistral model side by side.