The Towards AI Writing Benchmark · LLM Leaderboard
Can an AI model write my YouTube scripts?
The real test is simple: could I read a model's script out loud on camera without cringing? That's what ToneBench checks. Every model writes the same real YouTube scripts, all 10 from my What's AI briefs, five drafts each, and every draft is scored on nine writing dimensions against my finished version. Three LLM judges from three model families (Anthropic, OpenAI, DeepSeek) score it blind, and code handles what a judge shouldn't eyeball, like spoken length. My own scripts sit on the board too, as the human baseline, so you can see how close the models get to the real thing.
Start with the board below and click any row for that model's full scorecard. If you want to know whether cranking up reasoning effort buys better writing, the thinking-levels page plots score against cost and time for each effort setting. Before you quote a number, the methodology covers the judges, the human baseline and the trust checks. The about page has the story behind it.
A script can look polished on the page and still sound nothing like the person who has to say it out loud. That's why natural spoken delivery and voice fit are built into the rubric. To be clear, though, the scores on this page come from LLM judges, not from me reading every draft.
The 10 scripts, the nine metrics and what each one rewards, how my reference scripts became the human baseline, and the three-judge panel and its cross-family checks. If you plan to quote a number from this board, read this first.
Most of these scores move together: a model that writes well on one metric tends to write well on the others. The useful part is where a model breaks that pattern. So each of the nine dimensions gets its own card here, and you can see who wins on voice, who writes the best hooks, and who keeps it tight. Switch between all variants, the best configuration per model, or the best model per family. Drag, scroll or use the arrows, and click a model to open its scorecard.
Straight answers about what ToneBench measures, who's behind it, and how far you should trust the numbers. The full setup is on the methodology page, and the story and the team are on the about page.
It's a writing benchmark built around one narrow, practical question: can an AI model write a YouTube script in my voice? Every model writes the same 10 real What's AI video scripts from my briefs, and each draft is scored against my finished version of that script on voice, writing craft, substance, structure, hooks, flow, length discipline and AI-writing "slop".
So it measures voice and craft fidelity. It says nothing about general intelligence, and that's on purpose.
Me, Louis-François Bouchard, co-founder & CTO of Towards AI and the person behind the "What's AI" YouTube channel. I build and maintain it together with the Towards AI editorial team, who reviewed every reference script. The longer story is on the about page.
Public leaderboards measure broad preference: which answer a crowd or a judge likes better, on average, across everything. That's useful! But it tells you very little about whether a model can hold one specific voice, follow a strict style guide, stay away from AI clichés and land on a target length for a real production brief.
That's the gap I needed to measure, so that's what ToneBench measures. And the scripts come from my own production briefs instead of a public dataset, which keeps memorization much less of a worry (more on contamination below).
Every model gets the same prompt, about 11,000 tokens: a style guide built from my voice and storytelling guides, the target length, the brief and its research packet. It writes each script five times, because a single draft can get lucky, and every draft goes to a panel of three LLM judges. The judges don't know which model wrote what, they use the identical rubric and prompt for every model, and they have to quote evidence from the script before scoring each metric.
Seven of the nine metrics (tone, craft, substance, flow, YouTube practices, hook and visual cues) are the plain average of the three judges. Length isn't judged at all: code counts the spoken words and compares them to the target. Slop is half computed (the EQ-Bench slop-score pipeline plus our own banned-phrase scan) and half judged. The nine scores are then combined with editorial weights into the overall score out of 100.
Elo sits on top of those scores. For each script, every pair of models is compared on its rubric score, the higher one wins, and pairs too close to call count as draws. So Elo is a relative ranking of the same rubric results rather than a separate side-by-side vote, and resampling the run scores (a bootstrap) gives each rating its 95% interval.
The full metric list, weights, judge details and verification checks are on the methodology page.
Yes, and we measured it. When we re-judged the same stored outputs with different families as the judge, every family scored its own models highest, just by different amounts. That's why the judged scores come from a panel of three judges from three model families: Claude Opus 5 (Anthropic), GPT-5.6 Sol (OpenAI) and DeepSeek V4.1 Flash.
All three model lines also sit on the board as contestants, which is exactly the case the panel is for: every judged metric averages all three seats, so no model is ever scored by its own family alone. A new model only lands on the board once all three judges have scored it. On top of that, we run blind side-by-side A/B checks in both orders and an order-independent Bradley-Terry refit. The current per-judge agreement numbers are on the methodology page.
Tone: does it sound like me, the anti-hype, first-person, told-not-reported voice. Craft: sentence-level writing quality, the read-aloud test. Substance: accuracy, depth and grounding in the research packet. Hook: whether the opening earns the next thirty seconds. YouTube: structure and retention practices, like a genuine subscribe call that isn't bolted on and an open loop that pays off. Flow: continuity and emotion, each paragraph causing the next. Slop: how free the script is of overused AI phrases and patterns, half from a computed scan and half from the judges. Length: spoken length against the target, counted by code rather than judged, with overshooting penalized harder than coming in short. Cues: the quality of the visual [SHOW:] directions.
Tone carries the most weight and cues the least. The exact weights sit in the legend under the table.
The gold ★ row is my own scripts, in the versions that went through the Towards AI team's review (AI engineers, a professional writer, YouTube reviewers and a fact reviewer) and a final editing pass. They're scored with the same rubric and the same judges as every model (how the human baseline is made), so they're the bar the channel works to: a high bar, but not a theoretical ceiling. Those scripts land in the low 90s, not a literal 100.
One detail matters when you compare: the reference is scored on its own, while models are scored with the reference as the top anchor. So a model above the line writes at that level on this rubric. It didn't beat my script line for line. Where the top of the board stands against it right now is in the key findings and the board itself.
To keep contamination out. If the exact briefs, research packets and gold scripts were public, future models could train on them, and the benchmark would quietly measure memorization instead of writing. The exact rubric wording stays private too, so nobody can tune a model to the rubric instead of the writing.
New scripts are normally benchmarked before their video goes out, and all benchmark data carries an industry-standard canary marker. That lowers the risk, but none of it makes exposure impossible.
New models are added as soon as practical after release. New scripts come a few times a year, usually benchmarked before the video is out.
Whenever the rubric or a judge changes, the entire board is re-judged under the new setup, so scores from different judging regimes never get mixed, and the evaluation version is recorded with every result. That's what happened in September 2026, when the DeepSeek seat moved from V4 Flash to V4.1 Flash: every draft was judged again.
Not necessarily, and I'd be careful here. Each model's 95% Elo interval shows how much its rating moves when the recorded run scores are resampled and the ranking is recalculated. When two ranges sit far apart, that's strong evidence, as clear as this board gets. When they overlap, the pair is too close to call on this board.
Neither case is proof, though: checking whether two separate intervals overlap does not test the Elo gap between them, and overlap doesn't establish a tie. The tasks stay fixed, so the ranges also say nothing about how the order would change on new tasks. Read the score breakdowns and the actual example scripts alongside the rank.
No. ToneBench measures one demanding, real-world skill: writing long-form scripts in one specific human voice under a strict style contract. That correlates with careful instruction-following and style control, but it's deliberately narrow, and a model can top this board while being mid-pack at code, math or agentic work. Use it as one signal, next to benchmarks for the skills you need.
Yes, please. Missing models, wrong numbers, disagreements with the methodology, or interest in a similar evaluation for your own brand voice: send feedback. And if your team needs help choosing or deploying models in production, that's the kind of work Towards AI does.