Towards AITowards AIToneBench

Thinking levels · Meta

Does Muse Spark 1.1 write better when it thinks harder?

It does, and here the board can actually tell. Muse Spark 1.1 (thinking) beats Muse Spark 1.1 by 124 Elo and the ranges don't overlap, which is strong evidence for the thinking version. It costs $0.040 per script instead of $0.030, and you'll barely notice the extra wait.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
high1493 ±49-81.0$0.0400.5 min3,137#100
default (no effort flag)1370 ±49-79.5$0.0300.4 min862#110

Which setting should you use?

Thinking on, no hesitation. You pay 1.3 times as much, still pennies per script, and the board is confident the draft gets better. Easy trade. The default run spends far fewer reasoning tokens, so on this model the flag really changes how much thinking goes into each draft. That said, I wouldn't start with Muse Spark 1.1 anymore: Muse Spark 1.3 scores higher even without thinking, for about the same price.

Other Meta ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Meta model side by side.