🔍 Read the full analysis: How The Open TTS Leaderboard Evaluates Multilingual Speech And Voice Cloning At Scale on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Hugging Face has introduced the Open TTS Leaderboard, which compares text-to-speech models using speech-recognition error rates, generation speed and speaker similarity. It adds listening comparisons, but the project says its automated scores do not replace human judgments of naturalness or preference.
Hugging Face has launched the Open TTS Leaderboard, a system for comparing text-to-speech models on speech accuracy, inference speed and speaker similarity. The project says the automated evaluations can be completed in hours, offering a faster way to compare models as its Hub held more than 8,000 TTS models on September 30, 2026; it cautions that the scores do not determine which voices sound most natural or which listeners prefer.
The leaderboard measures accuracy by generating speech from prompts, transcribing that audio with Qwen3 automatic speech recognition and comparing the transcript with the original text. It reports word error rate and character error rate. The default ranking uses macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval; users can choose other languages and datasets.
For speed, the evaluation reports offline generation performance as inverse real-time factor on an Nvidia H200 GPU. It also measures streaming responsiveness using time-to-first-audio on an H200 and on a CPU. A separate voice-cloning view reports speaker similarity by comparing WavLM embeddings from generated speech and reference audio.
Hugging Face identifies k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models. For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among the leading systems. Those are results on the leaderboard’s chosen measures, not a broad determination of which products sound best.
A Faster Way to Compare Speech Models
The leaderboard addresses a practical problem for developers and researchers: text-to-speech models are released quickly, while comparisons based on listener votes can take substantially longer. Hugging Face says its objective evaluations can run in a couple of hours, compared with weeks for arena voting. Repeatable measurements may help people narrow a large field of models before testing audio themselves.
The distinct measures also help expose tradeoffs. A system with low transcription error may not be the fastest, and voice similarity does not tell users whether a model sounds expressive or natural. The streaming measurement may be useful to teams building voice agents, where time before playback begins affects how responsive an interaction feels. The scores can guide model selection, but they do not settle it.
Hugging Face says open-weight models are underrepresented in some existing rankings. It counted 16 open-weight models among 92 on Artificial Analysis as of September 30, 2026, and reported a similar skew on Voice Arena. The company attributes part of that imbalance to the work required to host open models and to commercial providers’ stronger incentives to seek placement. That is Hugging Face’s explanation, rather than an independently established cause.
multilingual text-to-speech device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Automated Scores and Listener Arenas
Text-to-speech systems turn written prompts into spoken audio. Existing comparison sites such as TTS Arena v2, Artificial Analysis and Voice Arena use paired listening comparisons: users hear outputs, cast votes and contribute to rankings often calculated using Elo scores and a Bradley–Terry model. These approaches directly capture preferences, but need a supply of votes and hosted model outputs.
The new leaderboard combines standardized datasets and automated measurements with a separate “Listen” tab. Users can select a language and dataset, compare generated audio, choose whether to evaluate voice cloning and submit preferences. Hugging Face says participants must sign in to reduce spam and bot submissions. The project describes the automated scores as a complement to, not a replacement for, human preference rankings.
The supplied announcement places the leaderboard in a fast-growing model ecosystem, reporting more than 8,000 TTS models on the Hugging Face Hub as of September 30, 2026. It presents faster, repeatable evaluation as a way to make comparisons easier across languages and operating needs, while leaving subjective judgments to listeners.
“The Open TTS Leaderboard does not replace human preference ranking.”
— Hugging Face, describing the leaderboard’s role
As an affiliate, we earn on qualifying purchases.
What the Scores Leave Out
The announcement does not provide the full model list, score tables, evaluation sample sizes or uncertainty ranges for individual results. It also does not establish how closely the automated metrics match listener judgments across languages, accents and speaking styles. Error rates depend on the speech-recognition system used, while speaker-similarity scores estimate identity preservation rather than naturalness, expressiveness or overall preference.
Coverage varies by language. The leaderboard uses character error rate for Chinese, Japanese and Korean. For languages beyond English and Chinese, Hugging Face says Seed TTS Eval has no audio and results come from CV3 Eval alone. The supplied material does not specify how often rankings will be refreshed, how model or dataset changes will be managed, or what threshold would be needed before community votes affect rankings.
Hugging Face says community votes may be incorporated as feedback accumulates, but it has not given a timetable or described how those votes would be combined with objective measurements. Rankings should therefore be interpreted as comparisons on specified tests and metrics, not as a universal ordering of voice quality.
As an affiliate, we earn on qualifying purchases.
Listening Tests and Ranking Updates
Users can try the leaderboard’s listening comparisons by selecting a language and dataset, then voting on generated audio. Hugging Face asks users to sign in to limit spam and bot activity. The project says it may add accumulated community preferences to the leaderboard, though no schedule or voting threshold has been announced.
For now, the available next step is to compare the published automated measures with the audio itself. Further clarity will depend on the project publishing more information about evaluation samples, score uncertainty, ranking refreshes and the role of community votes. Until then, the leaderboard offers a faster screening tool, with listener judgment still needed for qualities its metrics do not measure.
real-time speech synthesis hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the Open TTS Leaderboard measure?
It compares speech accuracy using word and character error rates, generation speed using offline and streaming measures, and voice-cloning similarity using WavLM embeddings.
Does a high leaderboard rank mean a model sounds best?
No. The measures do not directly assess naturalness, expressiveness or listener preference. Hugging Face says users should also listen to the audio.
How does the leaderboard evaluate speech accuracy?
It transcribes generated speech with Qwen3 automatic speech recognition and compares the transcript with the original prompt. Its default ranking uses macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval.
Can users contribute preferences?
Yes. Users can compare outputs and vote in the “Listen” tab, with sign-in required to help limit spam and bot submissions. Hugging Face says votes may be incorporated later but has not specified when or how.
When will the rankings be updated?
The announcement does not state a ranking update schedule or explain how changes to models and evaluation data will be handled.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
