The TTS evaluation problem has gotten embarrassing. The Hugging Face Hub hosts more than 8,000 text-to-speech models as of September 2026, but standardized evaluation infrastructure has lagged so far behind that open-source models remain systematically underrepresented on the arena-style leaderboards the community actually looks at. Hugging Face's new Open TTS Leaderboard is their attempt to fix that gap with objective metrics that can evaluate a model in hours instead of the weeks required to collect human votes.
This matters because the current gold-standard evaluation method—human preference scoring via arena comparisons—simply cannot scale. As of late September 2026, only 16 of 92 models on Artificial Analysis are open-weights, with similar skew on Voice Arena. The bottleneck is practical: adding a proprietary API model requires little more than an API key, whereas hosting an open model for user evaluation requires infrastructure and ongoing serving costs that arena operators must bear themselves.
What the Leaderboard Actually Measures
The Open TTS Leaderboard evaluates models on three complementary dimensions, none of which directly replace human preference but all of which provide fast, reproducible proxies:
Intelligibility via word error rate (WER) and character error rate (CER), computed by transcribing generated audio with Qwen3-ASR (the top-ranking open model on the Open ASR Leaderboard) and comparing it to the original prompt text. This doesn't measure naturalness or expressiveness, but it does tell you whether the model is actually saying what you asked it to say.
Speed via two metrics: inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, and time-to-first-audio (TTFA) for streaming latency at batch size 1 on both H200 GPU and CPU. TTFA quantifies how long a user waits after sending a prompt before playback can begin—critical for voice agents and interactive applications.
Speaker similarity (SIM) by computing cosine similarity between WavLM speaker embeddings of the generated audio and the reference clip. This estimates voice identity preservation in cloning scenarios, though it doesn't capture subjective qualities like warmth or character.
By default, models are ranked by macro-average WER across English splits of Seed TTS Eval and CV3 Eval (zero-shot). hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro lead the English WER rankings. Pareto plots visualize which models strike favorable tradeoffs between WER, batched inference speed, and model size.
Multilingual Reality Check
English performance is not a proxy for other languages, and the leaderboard design acknowledges this explicitly. You can toggle multiple languages to rank models on multilingual capability. Seed TTS Eval includes audio for English and Chinese; other languages draw scores from CV3 Eval (zero-shot). Chinese, Japanese, and Korean report character error rate (CER) instead of WER, and the cross-language "Average WER" is a macro-average.
k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 emerge as strong multilingual models. The voice-cloning toggle filters to models supporting that functionality on selected languages and adds a SIM column plus Pareto plots for the tradeoff between speaker similarity, inference speed, and size.
Interestingly, some models—bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 among them—show improved average WER under voice cloning, meaning they perform better when given a reference audio clip. That's a useful signal about which models benefit from few-shot conditioning.
The Listen Tab: Subjective Evaluation, Scaled
Numbers tell part of the story, but the leaderboard includes a "Listen" tab where you can compare generated outputs directly. Pick a language and dataset, optionally filter by voice-cloning capability, and either select specific models or hear outputs from a random selection. You can vote on which samples you prefer.
This fills a real gap: a centralized space to explore outputs from many models on the same prompts. As the team collects more community votes, they may incorporate that preference data into the rankings. The interface asks users to log in with Hugging Face accounts to filter spam and bots, which is the correct lightweight friction to add.
Streaming Performance on GPU and CPU
The "Streaming" tab ranks models by TTFA—time-to-first-audio—on 50 English prompts from CV3-Eval, batch size 1, using each model's default voice. For streaming models, TTFA measures the time until the first audio chunk arrives; for non-streaming models, it's the time until the complete utterance is generated, since playback cannot start earlier. The leaderboard drops the first three runs as warm-up and reports the median.
Default view shows H200 GPU performance, with a growing set of CPU results. kyutai/pocket-tts performs well for streaming on both GPU and CPU, which makes it especially interesting for edge deployment and low-latency interactive applications.
What This Leaderboard Is Not
The team is explicit about scope: objective metrics do not replace human preference ranking. ASR-based WER proxies intelligibility; speaker similarity estimates identity preservation. Neither measures naturalness, expressiveness, or listener preference directly. But fast, reproducible benchmarks can inform which models voting-based leaderboards should bother evaluating in the first place.
Arena-style evaluation also suffers from voter consistency problems. No arena can ensure the same voters with the same criteria for "better" evaluate models over time. Individual preferences drift. The Open TTS Leaderboard sidesteps this by fixing the evaluation criteria—imperfect proxies, but stable ones.
The Open-Source Evaluation Scripts Are Coming
Hugging Face plans to open-source the evaluation scripts, similar to the Open ASR Leaderboard repository, so the community can provide feedback and suggestions via GitHub issues and pull requests. This is the correct move: leaderboards ossify quickly if they're not designed as living infrastructure shaped by the people who actually use them.
The team explicitly frames the leaderboard as community-driven and wants input on which datasets, models, and metrics to add. They've focused initially on open-source models (to surface neglected work) and multilingual evaluation (because English is not a suitable proxy for other languages).
Why This Matters Beyond TTS
The Open TTS Leaderboard is tackling a problem that will recur in every sufficiently fast-moving modality: when model releases outpace evaluation infrastructure by an order of magnitude, open models get systematically excluded from the benchmarks that shape community perception and adoption. Proprietary models with marketing budgets and easy API access dominate leaderboards not because they're better, but because they're easier to add.
Objective metrics are imperfect, but they scale. Human preference is the ultimate ground truth, but it's expensive and slow. The smart move is using fast objective benchmarks as a first-pass filter and funneling the most promising models into deeper human evaluation. This leaderboard demonstrates that workflow.
The streaming performance tab is especially forward-looking. As voice agents become more common, TTFA on CPU will matter as much as raw WER. Models that sound great but take three seconds to start talking are not viable for interactive use cases. Surfacing that dimension now, while the field is still iterating rapidly, increases the chance that model developers will optimize for it.
The Bottom Line
Hugging Face's Open TTS Leaderboard is the evaluation infrastructure the open TTS ecosystem has needed for at least a year. It won't replace human preference ranking, and it's not trying to. But by making it possible to evaluate a model in hours instead of weeks, it levels the playing field for the 8,000+ open models that couldn't afford the infrastructure overhead or marketing push to get onto proprietary-dominated arenas.
The fact that multilingual support and voice-cloning evaluation are baked into the design from day one is a strong signal about priorities. The Listen tab turns the leaderboard into a discovery tool, not just a ranking. And the streaming performance focus shows the team is thinking about the deployment contexts that actually matter for voice applications.
If you're building with TTS models or tracking the open-source audio landscape, this leaderboard is now a bookmark-and-check-weekly resource. And if you have opinions about which metrics or datasets should be added, the team wants to hear them—this is infrastructure designed to evolve with the community, not ossify into a static benchmark suite.