Towards AITowards AIToneBench

Grok 4.6 (high) vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 (high) leads overall, 88.1 to 83.3.

Grok 4.6 (high)
#18 Elo 2144 · 88.1/100
Qwen3.8 Max
#41 Elo 1799 · 83.3/100
Cost / script
$0.115 vs $0.150
Human baseline
90.3 Grok 4.6 (high) falls below it · Qwen3.8 Max falls below it

The verdict

Grok 4.6 wins this one cleanly. It sits seventeenth on the board at 2144.5 Elo; Qwen3.8 Max is #42 at 1798.7, a gap of 345.8 points, and the confidence intervals do not overlap, so the gap is real. The overall scores say the same thing: 88.14 against 83.35, a difference of 4.79 points. The contrast is sharpest where scripts live or die. Grok hooks at 90.03 to Qwen's 83.95, and holds continuity and emotion at 86.27 where Qwen drops to 78.21. Qwen's one clear win is length adherence, 90.63 to 88.26, and it edges the slop metric 91.96 to 91.84. Consistency matters too: Qwen's overall swings with a 5.43 standard deviation against Grok's 1.43, so its average hides rougher outings. Then the unusual part: the better model is also the cheaper one. Grok runs about twelve cents per script, Qwen about fifteen cents, both exact figures. Neither is open weights, so there is no philosophical tiebreaker. This one is not close.

Pick Grok 4.6 (high) if you want the stronger script on nearly every metric, especially hooks and continuity, and you would rather pay less for it.
Pick Qwen3.8 Max if strict length adherence is the metric you care about most and you can accept wider swings in overall quality.

Metric by metric

Blue bars: Grok 4.6 (high). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
88.8
81.8
Writing Craft & Clarity13% weight
88.0
83.4
Substance, Accuracy & Value15% weight
88.0
84.6
Continuity & Emotion14% weight
86.3
78.2
YouTube Best Practices12% weight
87.1
80.8
Hook Strength10% weight
90.0
84.0
Length Adherence8% weight
88.3
90.6
Slop Score (EQ-Bench + ours)5% weight
91.8
92.0
Visual Cue Quality4% weight
86.4
84.8

Everything else that differs

Grok 4.6 (high)Qwen3.8 Max
Overall / 10088.183.3
Writing Elo21441799
Run-to-run spread (± overall std)1.4305.430
Cost per script (USD)0.1150.150
Avg latency (s)199.2395.6
Open weightsNoNo

Full scorecards: Grok 4.6 (high) · Qwen3.8 Max. How scoring works: methodology.

← All comparisons