AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How The Open TTS Leaderboard Evaluates Multilingual Speech And Voice Cloning At Scale on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face has introduced the Open TTS Leaderboard, which compares text-to-speech models using speech-recognition error rates, generation speed and speaker similarity. It adds listening comparisons, but the project says its automated scores do not replace human judgments of naturalness or preference.

Hugging Face has launched the Open TTS Leaderboard, a system for comparing text-to-speech models on speech accuracy, inference speed and speaker similarity. The project says the automated evaluations can be completed in hours, offering a faster way to compare models as its Hub held more than 8,000 TTS models on September 30, 2026; it cautions that the scores do not determine which voices sound most natural or which listeners prefer.

The leaderboard measures accuracy by generating speech from prompts, transcribing that audio with Qwen3 automatic speech recognition and comparing the transcript with the original text. It reports word error rate and character error rate. The default ranking uses macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval; users can choose other languages and datasets.

For speed, the evaluation reports offline generation performance as inverse real-time factor on an Nvidia H200 GPU. It also measures streaming responsiveness using time-to-first-audio on an H200 and on a CPU. A separate voice-cloning view reports speaker similarity by comparing WavLM embeddings from generated speech and reference audio.

Hugging Face identifies k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models. For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among the leading systems. Those are results on the leaderboard’s chosen measures, not a broad determination of which products sound best.

At a glance
announcementWhen: Announced September 2026
The developmentHugging Face has launched an open leaderboard for comparing text-to-speech models across accuracy, speed and voice-cloning measures.
At a glance
announcementWhen: Announced in material dated September 3…
The developmentHugging Face has launched an open leaderboard that evaluates text-to-speech models across multiple languages using objective performance metrics and offers audio comparisons for community feedback.

A Faster Way to Compare Speech Models

The leaderboard addresses a practical problem for developers and researchers: text-to-speech models are released quickly, while comparisons based on listener votes can take substantially longer. Hugging Face says its objective evaluations can run in a couple of hours, compared with weeks for arena voting. Repeatable measurements may help people narrow a large field of models before testing audio themselves.

The distinct measures also help expose tradeoffs. A system with low transcription error may not be the fastest, and voice similarity does not tell users whether a model sounds expressive or natural. The streaming measurement may be useful to teams building voice agents, where time before playback begins affects how responsive an interaction feels. The scores can guide model selection, but they do not settle it.

Hugging Face says open-weight models are underrepresented in some existing rankings. It counted 16 open-weight models among 92 on Artificial Analysis as of September 30, 2026, and reported a similar skew on Voice Arena. The company attributes part of that imbalance to the work required to host open models and to commercial providers’ stronger incentives to seek placement. That is Hugging Face’s explanation, rather than an independently established cause.

Amazon

multilingual text-to-speech device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Automated Scores and Listener Arenas

Text-to-speech systems turn written prompts into spoken audio. Existing comparison sites such as TTS Arena v2, Artificial Analysis and Voice Arena use paired listening comparisons: users hear outputs, cast votes and contribute to rankings often calculated using Elo scores and a Bradley–Terry model. These approaches directly capture preferences, but need a supply of votes and hosted model outputs.

The new leaderboard combines standardized datasets and automated measurements with a separate “Listen” tab. Users can select a language and dataset, compare generated audio, choose whether to evaluate voice cloning and submit preferences. Hugging Face says participants must sign in to reduce spam and bot submissions. The project describes the automated scores as a complement to, not a replacement for, human preference rankings.

The supplied announcement places the leaderboard in a fast-growing model ecosystem, reporting more than 8,000 TTS models on the Hugging Face Hub as of September 30, 2026. It presents faster, repeatable evaluation as a way to make comparisons easier across languages and operating needs, while leaving subjective judgments to listeners.

“The Open TTS Leaderboard does not replace human preference ranking.”

— Hugging Face, describing the leaderboard’s role

Amazon

voice cloning microphone

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Leave Out

The announcement does not provide the full model list, score tables, evaluation sample sizes or uncertainty ranges for individual results. It also does not establish how closely the automated metrics match listener judgments across languages, accents and speaking styles. Error rates depend on the speech-recognition system used, while speaker-similarity scores estimate identity preservation rather than naturalness, expressiveness or overall preference.

Coverage varies by language. The leaderboard uses character error rate for Chinese, Japanese and Korean. For languages beyond English and Chinese, Hugging Face says Seed TTS Eval has no audio and results come from CV3 Eval alone. The supplied material does not specify how often rankings will be refreshed, how model or dataset changes will be managed, or what threshold would be needed before community votes affect rankings.

Hugging Face says community votes may be incorporated as feedback accumulates, but it has not given a timetable or described how those votes would be combined with objective measurements. Rankings should therefore be interpreted as comparisons on specified tests and metrics, not as a universal ordering of voice quality.

Amazon

AI speech recognition headset

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Listening Tests and Ranking Updates

Users can try the leaderboard’s listening comparisons by selecting a language and dataset, then voting on generated audio. Hugging Face asks users to sign in to limit spam and bot activity. The project says it may add accumulated community preferences to the leaderboard, though no schedule or voting threshold has been announced.

For now, the available next step is to compare the published automated measures with the audio itself. Further clarity will depend on the project publishing more information about evaluation samples, score uncertainty, ranking refreshes and the role of community votes. Until then, the leaderboard offers a faster screening tool, with listener judgment still needed for qualities its metrics do not measure.

Amazon

real-time speech synthesis hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the Open TTS Leaderboard measure?

It compares speech accuracy using word and character error rates, generation speed using offline and streaming measures, and voice-cloning similarity using WavLM embeddings.

Does a high leaderboard rank mean a model sounds best?

No. The measures do not directly assess naturalness, expressiveness or listener preference. Hugging Face says users should also listen to the audio.

How does the leaderboard evaluate speech accuracy?

It transcribes generated speech with Qwen3 automatic speech recognition and compares the transcript with the original prompt. Its default ranking uses macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval.

Can users contribute preferences?

Yes. Users can compare outputs and vote in the “Listen” tab, with sign-in required to help limit spam and bot submissions. Hugging Face says votes may be incorporated later but has not specified when or how.

When will the rankings be updated?

The announcement does not state a ranking update schedule or explain how changes to models and evaluation data will be handled.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AutoSynthData: Generating Training Data For Enterprise Agents

ServiceNow’s AutoSynthData uses agent failures and teacher-model successes to generate and validate enterprise training tasks.

AI 2040: Plan A

The global initiative AI 2040: Plan A was announced today, outlining a comprehensive vision for artificial intelligence development through 2040.

Using ChatGPT Rank Tools To Improve AEO And GEO Performance

Exploring how ChatGPT rank monitoring tools can improve brand visibility in AI-driven search environments for mid-market and enterprise brands.

AI-Washed: When ‘Productivity’ Becomes the Press Release for Cuts You Couldn’t Justify

Tech giants like Meta and Microsoft announced 20,000 layoffs in April 2026, framing it as AI-driven efficiency. New data reveals the real story behind these cuts.