Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 81.7.
Grok 4.5 at #30 against Qwen3.7 Max at #50, 85.8 overall to 81.73. Under four points, and this one is genuinely interesting because the two models fail in opposite directions. Grok wins voice, craft, substance, and slop resistance: 87.07 against 81.15 on tone, 85.58 against 77.84 on anti-slop. Qwen wins length adherence and it isn't close, 85.79 against Grok's 78.67. So Grok writes better sentences that come out the wrong length, and Qwen hits the word count with flatter prose. Which one you want depends entirely on which of those you'd rather fix. Trimming a good draft to length is usually easier than injecting voice into a correctly-sized one, which is roughly why Grok ranks seventeen places higher. Grok is also cheaper at under four cents against five and a half, and four times faster at 42 seconds. On price, speed and quality it's the better pick, with length as the one real caveat.
Pick Grok 4.5 if you want stronger writing, faster and cheaper, and you don't mind trimming to length.
Pick Qwen3.7 Max if hitting a target word count without editing is what you're optimising for.
Blue bars: Grok 4.5. Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.
| Grok 4.5 | Qwen3.7 Max (default) | |
|---|---|---|
| Overall / 100 | 85.8 | 81.7 |
| Writing Elo | 2042 | 1699 |
| Run-to-run spread (± overall std) | 2.810 | 3.060 |
| Cost per script (USD) | 0.038 | 0.054 |
| Avg latency (s) | 41.7 | 156.8 |
| Open weights | No | No |
Full scorecards: Grok 4.5 · Qwen3.7 Max (default). How scoring works: methodology.
← All comparisons