Towards AITowards AIToneBench

Grok 4.6 vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 leads overall, 86.6 to 81.1.

Grok 4.6
#26 Elo 2035 · 86.6/100
Qwen3.8 Max
#46 Elo 1714 · 81.1/100
Cost / script
$0.204 vs $0.164
Human baseline
90.2 Grok 4.6 falls below it · Qwen3.8 Max falls below it

The verdict

Grok 4.6 wins this one comfortably. It sits at #26 with 2035.3 Elo against Qwen3.8 Max at #46 and 1713.9, and the confidence intervals don't come close to touching, so the 321.4-point gap is real. Overall it's 86.61 to 81.06, a 5.55-point spread, and the damage lands where a script lives or dies: continuity and emotion, 85.78 vs 76.22, YouTube best practices, 86.54 vs 77.76, and tone and voice, 87.86 vs 80.65. Qwen also swings much harder from script to script, a 7.89 standard deviation against Grok's 2.53. It takes exactly one metric back: hook strength, 83.39 to Grok's 80.81. The records say the same thing, Grok at 941 wins and 99 losses against Qwen's 588 and 258. Normally the loser gets a price argument, but there isn't much of one here: about sixteen cents per script against roughly twenty cents, both closed weights. Four cents doesn't buy back a 5.55-point gap.

Pick Grok 4.6 if you want the clearly stronger script on nearly every metric — voice, continuity, structure — delivered with far more consistency at roughly twenty cents each.
Pick Qwen3.8 Max if you want slightly stronger hooks and the cheaper draft at about sixteen cents per script, and you're prepared to rewrite around a 5.55-point overall gap.

Metric by metric

Blue bars: Grok 4.6. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.9
80.7
Writing Craft & Clarity13% weight
86.5
81.0
Substance, Accuracy & Value15% weight
87.5
81.3
Continuity & Emotion14% weight
85.8
76.2
YouTube Best Practices12% weight
86.5
77.8
Hook Strength10% weight
80.8
83.4
Length Adherence8% weight
88.0
86.3
Slop Score (EQ-Bench + ours)5% weight
91.2
90.4
Visual Cue Quality4% weight
86.6
80.6

Everything else that differs

Grok 4.6Qwen3.8 Max
Overall / 10086.681.1
Writing Elo20351714
Run-to-run spread (± overall std)2.5307.890
Cost per script (USD)0.2040.164
Avg latency (s)243.1429.0
Open weightsNoNo

Full scorecards: Grok 4.6 · Qwen3.8 Max. How scoring works: methodology.

← All comparisons