Towards AITowards AIToneBench

Claude Opus 5 (max) vs Grok 4.6 (high)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.4 to 88.1.

Claude Opus 5 (max)
#1 Elo 2358 · 90.4/100
Grok 4.6 (high)
#18 Elo 2144 · 88.1/100
Cost / script
$0.127 vs $0.115
Human baseline
90.3 Claude Opus 5 (max) exceeds it · Grok 4.6 (high) falls below it

The verdict

No contest on the board. Opus 5 at max effort sits first at 2358.5 Elo; Grok 4.6 high is seventeenth at 2144.5, and the confidence intervals, 2312 to 2409 against 2103 to 2181, are not close to touching. The overall scores tell the same story: 90.37 against 88.14, a 2.23-point gap, which is large for this leaderboard. The separation is widest on continuity and emotion, 90.17 to 86.27, and Opus also leads on cue quality, 88.24 to 86.43. Grok claims exactly one metric: length adherence, 88.26 to 87.2. The usual escape hatch, price, is shut here. Opus costs about thirteen cents per script (estimated) and Grok about twelve cents (exact), so Grok is not meaningfully cheaper. Neither is open weights. A rare case where the better writer costs about a penny more, and we would pay it.

Pick Claude Opus 5 (max effort) if you want the strongest writer on the board, especially for continuity and cue quality, and a roughly one-penny premium per script is irrelevant.
Pick Grok 4.6 (high) if you value exact, predictable billing and slightly tighter length adherence, and can accept a clear step down in overall writing quality.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: Grok 4.6 (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
90.9
88.8
Writing Craft & Clarity13% weight
91.0
88.0
Substance, Accuracy & Value15% weight
90.3
88.0
Continuity & Emotion14% weight
90.2
86.3
YouTube Best Practices12% weight
89.6
87.1
Hook Strength10% weight
91.9
90.0
Length Adherence8% weight
87.2
88.3
Slop Score (EQ-Bench + ours)5% weight
93.0
91.8
Visual Cue Quality4% weight
88.2
86.4

Everything else that differs

Claude Opus 5 (max)Grok 4.6 (high)
Overall / 10090.488.1
Writing Elo23582144
Run-to-run spread (± overall std)1.5301.430
Cost per script (USD)0.1270.115
Avg latency (s)244.2199.2
Open weightsNoNo

Full scorecards: Claude Opus 5 (max) · Grok 4.6 (high). How scoring works: methodology.

← All comparisons