Towards AITowards AIToneBench

Grok 4.5 vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 83.4.

Grok 4.5
#30 Elo 2042 · 85.8/100
Qwen3.8 Max
#38 Elo 1866 · 83.4/100
Cost / script
$0.038 vs $0.150
Human baseline
92.7 Grok 4.5 falls below it · Qwen3.8 Max falls below it

The verdict

Grok 4.5 at #30 and Qwen3.8 Max at #38, 85.80 overall to 83.44. Only one and a half points, and the two fail in opposite directions, which makes this a real choice rather than a ranking. Grok wins voice 87.07 to 81.93, continuity 83.50 to 78.08, and hooks 87.96 to 83.50. Qwen wins length adherence and it is not close: 92.90 against 78.67, over sixteen points. So Grok writes the better script at the wrong length; Qwen hits the word count with flatter prose. Which you prefer depends on which you would rather fix, and trimming a good draft is usually easier than injecting voice into a correctly-sized one. Grok is also four times cheaper at under four cents against 15, and ten times faster at 37 seconds against 397. On price, speed and voice it is the better pick, with length as the one real caveat.

Pick Grok 4.5 for better writing, much cheaper and much faster, if you can trim to length.
Pick Qwen3.8 Max if hitting an exact word count without editing is the thing you are optimising for.

Metric by metric

Blue bars: Grok 4.5. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.1
81.9
Writing Craft & Clarity13% weight
86.6
83.4
Substance, Accuracy & Value15% weight
87.4
84.4
Continuity & Emotion14% weight
83.5
78.1
YouTube Best Practices12% weight
84.4
80.4
Hook Strength10% weight
88.0
83.5
Length Adherence8% weight
78.7
92.9
Slop Score (EQ-Bench + ours)5% weight
90.2
92.0
Visual Cue Quality4% weight
86.5
85.2

Everything else that differs

Grok 4.5Qwen3.8 Max
Overall / 10085.883.4
Writing Elo20421866
Run-to-run spread (± overall std)2.8105.620
Cost per script (USD)0.0380.150
Avg latency (s)41.7397.4
Open weightsNoNo

Full scorecards: Grok 4.5 · Qwen3.8 Max. How scoring works: methodology.

← All comparisons