Where this benchmark comes from, and who is behind it.
ToneBench grew out of the scripts we write for our YouTube channel, and the broader work we do at Towards AI helping people and companies actually build with AI. We pick writing models for this exact job every week, so we wanted a measurement we could trust more than a generic leaderboard.
One of the largest, most trusted AI learning communities: a widely-read Medium publication and newsletter, and the team behind the book Building LLMs for Production. We're engineers and researchers who care about production-grade AI over fragile demos. More about us →
The benchmark is built and maintained by Louis-François Bouchard, co-founder & CTO of Towards AI and creator of the "What's AI" channel, with the Towards AI editorial team. How the scores are produced is on the methodology page.
Hands-on courses to become an AI engineer who actually builds and ships, from Master AI Engineering to Agent Engineering and more. You finish with a deployed product and a certified portfolio, not just notebooks.
We help teams adopt AI for real: enablement and upskilling for your engineers, plus custom AI development and value creation when you'd rather we build and advise alongside you.
One of the largest and most trusted AI learning communities: a widely-read Medium publication and newsletter, the team behind Building LLMs for Production, hands-on AI-engineering courses, and enterprise AI work. ToneBench is our writing benchmark.
Louis-François Bouchard, co-founder and CTO of Towards AI. Practical, no-hype AI-engineering videos, and where this benchmark's reference scripts come from. Subscribe here.
Through hands-on courses that end with a deployed product and a certified portfolio you can show. Start at the Academy or read Building LLMs for Production.
Yes: enablement and upskilling for engineering teams, plus custom AI development when you would rather we build alongside you. See enterprise enablement, value creation, and case studies.
Every current-ranked model writes the same 9 real scripts, five times each, scored blind against our finished versions across nine dimensions by a panel of three LLM judges from three different families. Details on the methodology page; ranking on the leaderboard.