Towards AITowards AIToneBench

Every Meta model, side by side

If you're hoping an open Meta model can write your scripts, this page is the honest answer. The Llama rows sit near the very bottom of the board, while the closed Muse Spark models are the only Meta options I'd consider, led by Muse Spark 1.3 (thinking) at #83. That's really two families under one logo, and it's why Llama 4 Scout, the last row here, trails the top by 1858 Elo.

Best writer
Muse Spark 1.3 (thinking)#83 · Elo 1629 · $0.041 per script
Best value (within 200 Elo of the best)
Muse Spark 1.3#88 · Elo 1600 · $0.034 per script
Cheapest
Llama 3.1 8B#178 · Elo -223 · $0.0007 per script

All 7 configurations

Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.

Board rankConfigurationEloOverallToneCraftSubstanceHookYouTubeFlowSlopLengthCuesCost / scriptTime / script
#83Muse Spark 1.3 (thinking)1629 ±5482.585.284.483.685.284.582.287.963.977.7$0.0411.3 min
#88Muse Spark 1.31600 ±5082.485.284.483.784.884.382.387.763.377.9$0.0341.1 min
#100Muse Spark 1.1 (thinking)1493 ±4981.083.681.582.084.882.579.685.262.984.0$0.0400.5 min
#110Muse Spark 1.11370 ±4979.583.581.481.085.080.678.684.153.382.0$0.0300.4 min
#169Llama 4 Maverick · open348 ±3260.559.667.665.768.653.960.973.930.763.8$0.00320.3 min
#178Llama 3.1 8B · open-223 ±3945.444.550.250.956.135.840.068.129.138.2$0.00070.1 min
#179Llama 4 Scout · open-229 ±3245.249.655.750.442.638.243.370.414.034.9$0.00240.3 min

Which one should you use?

I'd use Muse Spark 1.3 (thinking). If you'd rather save a bit, Muse Spark 1.3 is the value pick, 29 Elo behind for $0.034 instead of $0.041, and to be honest the board can't separate those two, so either one is fine. Both score lowest on length adherence, by a wide margin, so check the word count before you publish. Muse Spark 1.1 is the older version, and I don't see a reason to pick it over Muse Spark 1.3 now. As for Llama 3.1 8B, it's a small, older model that costs next to nothing, so this isn't really a knock on it. But at #178 it's not a script writer, and I wouldn't use any Llama row for this.

Thinking levels, model by model

Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.

Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.

← All comparisons