Thinking levels · Alibaba (Qwen)
No, at least not in anything the board measured. The high effort flag and the default end up 9 Elo apart with overlapping ranges. They also land at about the same cost and the same wait, and even reason about as much, so on this model the flag is close to a no-op.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Reasoning tokens | Board rank |
|---|---|---|---|---|---|---|---|
| high | 1447 ±40 | - | 80.2 | $0.092 | 1.7 min | 6,037 | #106 |
| default (no effort flag) | 1456 ±37 | - | 80.4 | $0.094 | 1.8 min | 6,312 | #103 |
The 95% Elo ranges overlap: Qwen3.7 Max (default) runs from 1414 to 1488, and Qwen3.7 Max (high) from 1406 to 1486.
Don't bother with the flag. There's no measurable gain from high, so there's nothing to lose by staying on default. The model you pick matters a lot more than the flag. Qwen3.7 Max sits at #103, while the newer Qwen3.8 Flash is far ahead at #79 for $0.013 per script, against $0.094 here. It's slower, though, at around 5.7 minutes per script against 1.8, so if speed is what keeps you on Qwen3.7 Max, that's a fair reason to stay.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Alibaba (Qwen) model side by side.