Towards AITowards AIToneBench

Grok 4.5 vs Qwen3.7 Max (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 81.5.

Grok 4.5
#30 Elo 2042 · 85.8/100
Qwen3.7 Max (high)
#55 Elo 1674 · 81.5/100
Cost / script
$0.038 vs $0.054
Human baseline
92.7 Grok 4.5 falls below it · Qwen3.7 Max (high) falls below it

The verdict

The rare matchup where one model wins on quality, price, and speed at the same time. Grok 4.5 scores 85.8 overall to Qwen3.7 Max's 81.5, costs about two cents a script against Qwen's estimated five, and returns a draft in under 30 seconds versus over two and a half minutes. The Elo intervals don't overlap either. Where the writing actually differs is tone. Grok sits at 87.1 on voice match, Qwen at 80.7, and that gap is wider than both standard deviations combined, so I treat it as real: Qwen reads assembled, not spoken. Its visual cues are also a weak spot here at 81.0. Credit where due, though. Qwen wins length adherence, 83.3 to Grok's 78.7, and Grok's length variance is no joke, so Qwen drafts come back the right size more often. That matters if trimming is the part you hate. But both are closed models, so Qwen can't play the open-weights card, and paying more for the weaker script is a tough sell. Grok, comfortably. Neither is anywhere near the human scripts yet, for the record.

Pick Grok 4.5 if you want the stronger voice at less than half the cost and a fraction of the wait.
Pick Qwen3.7 Max if hitting target length on the first pass is your biggest pain point.

Metric by metric

Blue bars: Grok 4.5. Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.1
80.9
Writing Craft & Clarity13% weight
86.6
81.5
Substance, Accuracy & Value15% weight
87.4
81.2
Continuity & Emotion14% weight
83.5
79.2
YouTube Best Practices12% weight
84.4
80.7
Hook Strength10% weight
88.0
85.0
Length Adherence8% weight
78.7
83.3
Slop Score (EQ-Bench + ours)5% weight
90.2
85.9
Visual Cue Quality4% weight
86.5
78.4

Everything else that differs

Grok 4.5Qwen3.7 Max (high)
Overall / 10085.881.5
Writing Elo20421674
Run-to-run spread (± overall std)2.8102.580
Cost per script (USD)0.0380.054
Avg latency (s)41.7157.2
Open weightsNoNo

Full scorecards: Grok 4.5 · Qwen3.7 Max (high). How scoring works: methodology.

← All comparisons