If you're hoping an open Meta model can write your scripts, this page is the honest answer. The Llama rows sit near the very bottom of the board, while the closed Muse Spark models are the only Meta options I'd consider, led by Muse Spark 1.3 (thinking) at #83. That's really two families under one logo, and it's why Llama 4 Scout, the last row here, trails the top by 1858 Elo.
Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.
| Board rank | Configuration | Elo | Overall | Tone | Craft | Substance | Hook | YouTube | Flow | Slop | Length | Cues | Cost / script | Time / script |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| #83 | Muse Spark 1.3 (thinking) | 1629 ±54 | 82.5 | 85.2 | 84.4 | 83.6 | 85.2 | 84.5 | 82.2 | 87.9 | 63.9 | 77.7 | $0.041 | 1.3 min |
| #88 | Muse Spark 1.3 | 1600 ±50 | 82.4 | 85.2 | 84.4 | 83.7 | 84.8 | 84.3 | 82.3 | 87.7 | 63.3 | 77.9 | $0.034 | 1.1 min |
| #100 | Muse Spark 1.1 (thinking) | 1493 ±49 | 81.0 | 83.6 | 81.5 | 82.0 | 84.8 | 82.5 | 79.6 | 85.2 | 62.9 | 84.0 | $0.040 | 0.5 min |
| #110 | Muse Spark 1.1 | 1370 ±49 | 79.5 | 83.5 | 81.4 | 81.0 | 85.0 | 80.6 | 78.6 | 84.1 | 53.3 | 82.0 | $0.030 | 0.4 min |
| #169 | Llama 4 Maverick · open | 348 ±32 | 60.5 | 59.6 | 67.6 | 65.7 | 68.6 | 53.9 | 60.9 | 73.9 | 30.7 | 63.8 | $0.0032 | 0.3 min |
| #178 | Llama 3.1 8B · open | -223 ±39 | 45.4 | 44.5 | 50.2 | 50.9 | 56.1 | 35.8 | 40.0 | 68.1 | 29.1 | 38.2 | $0.0007 | 0.1 min |
| #179 | Llama 4 Scout · open | -229 ±32 | 45.2 | 49.6 | 55.7 | 50.4 | 42.6 | 38.2 | 43.3 | 70.4 | 14.0 | 34.9 | $0.0024 | 0.3 min |
I'd use Muse Spark 1.3 (thinking). If you'd rather save a bit, Muse Spark 1.3 is the value pick, 29 Elo behind for $0.034 instead of $0.041, and to be honest the board can't separate those two, so either one is fine. Both score lowest on length adherence, by a wide margin, so check the word count before you publish. Muse Spark 1.1 is the older version, and I don't see a reason to pick it over Muse Spark 1.3 now. As for Llama 3.1 8B, it's a small, older model that costs next to nothing, so this isn't really a knock on it. But at #178 it's not a script writer, and I wouldn't use any Llama row for this.
Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.
Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.
← All comparisons