Towards AITowards AIToneBench

GLM-5.3 vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 86.8 to 81.1.

GLM-5.3
#24 Elo 2042 · 86.8/100
Qwen3.8 Max
#46 Elo 1714 · 81.1/100
Cost / script
$0.109 vs $0.164
Human baseline
90.2 GLM-5.3 falls below it · Qwen3.8 Max falls below it

The verdict

GLM-5.3 wins this one outright, and there is no angle where Qwen3.8 Max claws it back. GLM sits at #24 with 2041.5 Elo against Qwen's #46 at 1713.9, a 327.6-point gap, and the confidence intervals aren't close to touching, so the ranking is real. Overall it's 86.83 to 81.06, a 5.77-point gap, and GLM sweeps every single metric. The widest split is continuity and emotion, 85.11 vs 76.22, which in practice means Qwen loses the thread of the story more often. Hook strength goes 89.54 to 83.39, tone and voice 87.89 to 80.65. Qwen only stays within about a point on the mechanical stuff: length adherence, 86.34 to GLM's 87.5, and visual cues, 80.59 to 81.43. The records say the same thing, GLM at 866 wins and 8 losses against Qwen's 588 and 258. Then the part that ends the debate: GLM is also cheaper, about eleven cents per script against about sixteen, faster per run, and it's open weights at 400B parameters while Qwen is a closed trillion-parameter model. Better, cheaper, and open. There is no trade to weigh here.

Pick GLM-5.3 if you want the stronger script on every metric, open weights at 400B parameters you can run yourself, and the lower bill at about eleven cents per draft.
Pick Qwen3.8 Max only if you're already committed to its hosted API and the two metrics where it nearly keeps pace, length adherence and visual cues, are the ones you care about most.

Metric by metric

Blue bars: GLM-5.3. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.9
80.7
Writing Craft & Clarity13% weight
87.2
81.0
Substance, Accuracy & Value15% weight
85.9
81.3
Continuity & Emotion14% weight
85.1
76.2
YouTube Best Practices12% weight
85.0
77.8
Hook Strength10% weight
89.5
83.4
Length Adherence8% weight
87.5
86.3
Slop Score (EQ-Bench + ours)5% weight
91.6
90.4
Visual Cue Quality4% weight
81.4
80.6

Everything else that differs

GLM-5.3Qwen3.8 Max
Overall / 10086.881.1
Writing Elo20421714
Run-to-run spread (± overall std)6.6007.890
Cost per script (USD)0.1090.164
Avg latency (s)301.9429.0
Open weightsYesNo

Full scorecards: GLM-5.3 · Qwen3.8 Max. How scoring works: methodology.

← All comparisons