Towards AITowards AIToneBench

Grok 4.5 vs Qwen3.7 Max (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 81.7.

Grok 4.5
#30 Elo 2042 · 85.8/100
Qwen3.7 Max (default)
#50 Elo 1699 · 81.7/100
Cost / script
$0.038 vs $0.054
Human baseline
92.7 Grok 4.5 falls below it · Qwen3.7 Max (default) falls below it

The verdict

Grok 4.5 at #30 against Qwen3.7 Max at #50, 85.8 overall to 81.73. Under four points, and this one is genuinely interesting because the two models fail in opposite directions. Grok wins voice, craft, substance, and slop resistance: 87.07 against 81.15 on tone, 85.58 against 77.84 on anti-slop. Qwen wins length adherence and it isn't close, 85.79 against Grok's 78.67. So Grok writes better sentences that come out the wrong length, and Qwen hits the word count with flatter prose. Which one you want depends entirely on which of those you'd rather fix. Trimming a good draft to length is usually easier than injecting voice into a correctly-sized one, which is roughly why Grok ranks seventeen places higher. Grok is also cheaper at under four cents against five and a half, and four times faster at 42 seconds. On price, speed and quality it's the better pick, with length as the one real caveat.

Pick Grok 4.5 if you want stronger writing, faster and cheaper, and you don't mind trimming to length.
Pick Qwen3.7 Max if hitting a target word count without editing is what you're optimising for.

Metric by metric

Blue bars: Grok 4.5. Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.1
81.2
Writing Craft & Clarity13% weight
86.6
81.7
Substance, Accuracy & Value15% weight
87.4
80.9
Continuity & Emotion14% weight
83.5
79.0
YouTube Best Practices12% weight
84.4
80.8
Hook Strength10% weight
88.0
84.2
Length Adherence8% weight
78.7
85.8
Slop Score (EQ-Bench + ours)5% weight
90.2
86.0
Visual Cue Quality4% weight
86.5
80.2

Everything else that differs

Grok 4.5Qwen3.7 Max (default)
Overall / 10085.881.7
Writing Elo20421699
Run-to-run spread (± overall std)2.8103.060
Cost per script (USD)0.0380.054
Avg latency (s)41.7156.8
Open weightsNoNo

Full scorecards: Grok 4.5 · Qwen3.7 Max (default). How scoring works: methodology.

← All comparisons