Who built it, why, and what you can do with it.
I'm Louis-François Bouchard, co-founder and CTO of Towards AI, and I host the "What's AI" YouTube channel. I built ToneBench because I wanted to know which models can actually write a script the way I do: the whole thing, with the hook, the structure, the anti-hype voice and the technical depth, good enough to air without edits. Generic leaderboards can't answer that, so the benchmark is built on the scripts behind 10 of my own What's AI videos, and my finished version of each one is the human reference every model is measured against.
If you came here with a practical question, here's where each answer lives:
We started as a Medium publication, grew into one of the bigger AI learning communities with a newsletter to match, and wrote the book Building LLMs for Production. We're engineers and researchers, and mostly we care about AI that actually ships. More about us →
I build and maintain the benchmark, with the Towards AI editorial team, who reviewed every reference script before it became the bar. You can find more of my work on my website, and how the scores are produced is on the methodology page.
Hands-on courses for becoming an AI engineer who actually builds and ships, from Master AI Engineering to Agent Engineering and more. You finish with a deployed product and a certified portfolio you can show.
We train engineering teams to build with AI, and when you'd rather have us build it with you, we do that too. The case studies show what that looked like for other companies.
The company I co-founded, and the team behind this benchmark. We teach AI engineering, publish, and build AI with companies; the sections above have the details, and the rest is on towardsai.net. ToneBench is our writing benchmark.
I do. I'm Louis-François Bouchard, co-founder and CTO of Towards AI. The channel is practical, no-hype AI engineering, and it's where this benchmark's reference scripts come from. Subscribe here if that's your kind of thing.
Start at the Academy. The Learn with us card above covers what the courses give you, plus the book.
Yes. That's the Work with us section above: team training, custom builds, and the case studies.
Every current-ranked model writes the same 10 real scripts, five times each. Every draft is scored blind against my finished version on nine dimensions: seven by a panel of three LLM judges from three model families, one (length) in code, and one (slop) half in code, half by the judges. The full details, including how far to trust the ranking, are on the methodology page, and the ranking itself is on the leaderboard.