Towards AITowards AIToneBench

DeepSeek V4 Pro 0813 (max) vs Muse Spark 1.3 (thinking)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Pro 0813 (max) leads overall, 84.0 to 82.6.

DeepSeek V4 Pro 0813 (max)
#40 Elo 1861 · 84.0/100
Muse Spark 1.3 (thinking)
#49 Elo 1699 · 82.6/100
Cost / script
$0.022 vs $0.041
Human baseline
90.2 DeepSeek V4 Pro 0813 (max) falls below it · Muse Spark 1.3 (thinking) falls below it

The verdict

DeepSeek V4 Pro 0813 (max) has the stronger current result: #40 at 1861.4 Elo and 84.04 overall, compared with Muse Spark 1.3 (thinking) at #49, 1699.2 Elo, and 82.63 overall. Their 95% Elo confidence intervals do not overlap, which supports the directional ordering (1794.00–1932.20 and 1653.10–1736.10). DeepSeek V4 Pro 0813 (max) has its clearest metric edges in Length Adherence (86.24 versus 63.88) and Visual Cue Quality (79.47 versus 75.93). Muse Spark 1.3 (thinking) counters on YouTube Best Practices (83.72 versus 81.13) and Continuity & Emotion (82.41 versus 80.77). At the measured run mix, DeepSeek V4 Pro 0813 (max) costs $0.022 per article versus $0.041 for Muse Spark 1.3 (thinking); DeepSeek V4 Pro 0813 (max) is the cheaper route. DeepSeek V4 Pro 0813 (max) is the open-weights option; Muse Spark 1.3 (thinking) is closed. On the current automated evidence, DeepSeek V4 Pro 0813 (max) is the stronger default; Muse Spark 1.3 (thinking) remains a defensible choice when its specific strengths, price, or deployment profile matter more than the headline rank.

Pick DeepSeek V4 Pro 0813 (max) when you prioritize the stronger current board result, length adherence, visual cue quality, open weights, and lower measured cost.
Pick Muse Spark 1.3 (thinking) when you prioritize youtube best practices, and continuity & emotion.

Metric by metric

Blue bars: DeepSeek V4 Pro 0813 (max). Orange bars: Muse Spark 1.3 (thinking). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
84.7
85.0
Writing Craft & Clarity13% weight
84.9
84.1
Substance, Accuracy & Value15% weight
84.7
84.9
Continuity & Emotion14% weight
80.8
82.4
YouTube Best Practices12% weight
81.1
83.7
Hook Strength10% weight
86.8
86.3
Length Adherence8% weight
86.2
63.9
Slop Score (EQ-Bench + ours)5% weight
88.0
89.1
Visual Cue Quality4% weight
79.5
75.9

Everything else that differs

DeepSeek V4 Pro 0813 (max)Muse Spark 1.3 (thinking)
Overall / 10084.082.6
Writing Elo18611699
Run-to-run spread (± overall std)4.3702.770
Cost per script (USD)0.0220.041
Avg latency (s)195.278.0
Open weightsYesNo

Full scorecards: DeepSeek V4 Pro 0813 (max) · Muse Spark 1.3 (thinking). How scoring works: methodology.

← All comparisons