The Towards AI Writing Benchmark · LLM Leaderboard
Can a model write in our voice?
Every model writes the same real YouTube scripts, and a panel of three LLM judges from three different families scores each draft blind against our own finished versions across nine writing dimensions. Read the methodology for what we measure, the human baseline, the judges and trust checks, and why the reference scripts stay private, or the about page for where this comes from.
The 9 scripts, the nine metrics and what each rewards, the human baseline, the judge and cross-family trust checks, and why the references stay private. It is all on the methodology page.
Slide through the nine writing dimensions. Each card ranks every model on that one metric, so you can see who wins on voice, who wins on hooks, who keeps it tight. Drag, scroll, or use the arrows. Click a model to open its scorecard.
Quick answers about what ToneBench measures, who is behind it, and how far to trust the numbers. The full setup lives on the methodology page; the story and the team are on the about page.
ToneBench is a writing benchmark that asks one narrow, practical question: can an AI model write a YouTube script in our editorial voice? Every model writes the same set of real video briefs, the same articles our team actually produced, and each script is scored against our finished reference version on voice, writing craft, substance, structure, hooks, flow, length discipline, and AI-writing "slop". It is a measure of voice and craft fidelity, not general intelligence.
It is built and maintained by Louis-François Bouchard, co-founder & CTO of Towards AI and creator of the "What's AI" YouTube channel, together with the Towards AI editorial team. More on the team and where this comes from is on the about page.
Public leaderboards measure broad preference: which answer a crowd or a judge likes better, on average, across everything. That is useful, but it says little about whether a model can hold a specific voice, follow a strict style guide, resist AI clichés, and hit a target length on a real production brief. ToneBench measures exactly that gap. And because the tasks are our own unpublished briefs, models can't have memorized the answers.
Every model receives the same inputs: our writing style guide, a task brief, and a research packet. Each model writes every script several times (multiple runs, averaged), and every script is scored blind by a panel of three LLM judges from three different families (Anthropic, OpenAI, DeepSeek) against a fixed rubric. Each judge must quote evidence from the script before scoring each metric, every judge uses the identical rubric and prompt for every model, and the published per-metric score is the unweighted mean of the three judgments. Per-metric scores are combined with editorial weights, and the Elo ranking comes from pairwise comparisons of those scores with bootstrap confidence intervals. The exact rubric wording and reference scripts are withheld to keep the benchmark hard to game. The full metric list, weights, judge details, and verification checks live on the methodology page.
Yes. We measured it, and it is exactly why the published score comes from a panel of three judges from three different families: Claude Opus 5 (Anthropic), GPT-5.6 Sol (OpenAI), and DeepSeek V4 Flash. When we re-judged the same stored outputs with different families as judge, every family scored its own models highest, by different amounts. In the panel, each judge scores every draft blind with the identical rubric, the published score is the unweighted mean of the three, and no model is ever scored by its own family alone. We also run blind side-by-side A/B checks in both orders and an order-independent Bradley-Terry refit. The current per-judge agreement numbers live on the methodology page.
Tone: does it sound like our voice and persona. Craft: sentence-level writing quality. Value: substance and factual grounding in the provided research. Hook: strength of the opening. YouTube: structure and retention practices for the format. Flow: continuity and emotional pacing. Slop: absence of over-used AI phrases and patterns. Length: discipline against the target spoken length. Cues: quality of the visual [SHOW:] directions.
The dotted-gold line is our own team-reviewed reference scripts scored under the same rubric (how the human baseline is made). Those scripts went through multiple expert reviewers and a final editorial pass, so they represent the quality bar the channel actually ships: a high bar, not a theoretical ceiling. The current model-versus-baseline relationship is regenerated in the key findings and board data.
To prevent contamination. If the exact briefs, research packets, and gold scripts were public, future models could train on them and the benchmark would quietly measure memorization instead of writing ability. New articles are benchmarked before their videos are published, the addition date is recorded, and all benchmark data carries an industry-standard canary marker.
New models are added as soon as practical after release. We watch the major provider APIs daily. New articles are added a few times per year, benchmarked pre-publication. When the rubric or judge version changes, the entire board is re-judged under the new setup so scores are never mixed across judging regimes; the evaluation version is recorded with every result.
Maybe. Every model carries a 95% confidence interval; when two intervals overlap, the board does not claim their order is settled. Treat them as a tie. Many widely separated models are distinguishable, but inside the tightly-packed top ten many neighbours are genuine near-ties. The interval, not the rank number, is the honest read.
No. ToneBench measures one demanding, real-world skill: writing long-form scripts in a specific human voice under a strict style contract. That correlates with careful instruction-following and style control, but it is deliberately narrow: a model can top this board and still be mid-pack at code, math, or agentic work. Use it as one signal, alongside benchmarks for the skills you care about.
Yes. Missing models, wrong numbers, methodology disagreements, or interest in a similar evaluation for your own brand voice: send feedback. If your team wants help choosing or deploying writing models in production, that is exactly what Towards AI does.