Thinking levels · Meta
Maybe a tiny bit, and not by enough to prove it. Muse Spark 1.3 (thinking) scores 29 Elo above Muse Spark 1.3, but the ranges overlap, so the board can't separate them with confidence. Luckily, little is riding on it: thinking costs $0.041 per script against $0.034, and the wait barely changes.
| Setting | Elo | Step | Overall | Cost / script | Time / script | Reasoning tokens | Board rank |
|---|---|---|---|---|---|---|---|
| high | 1629 ±54 | - | 82.5 | $0.041 | 1.3 min | 3,355 | #83 |
| default (no effort flag) | 1600 ±50 | - | 82.4 | $0.034 | 1.1 min | 1,865 | #88 |
The 95% Elo ranges overlap: Muse Spark 1.3 runs from 1546 to 1645, and Muse Spark 1.3 (thinking) from 1568 to 1676.
I'd turn thinking on. It costs almost nothing extra, and whatever small lean the board shows points that way. Just don't expect it to fix what holds this model back. Both settings score about the same on every metric, including length adherence, which is Muse Spark's weakest by far. If length is your problem, thinking harder won't solve it, and you'll still be fixing that by hand.
Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Meta model side by side.