What this is

On Sept 30, Hugging Face launched the Open TTS Leaderboard, pulling more than 8,000 text-to-speech (TTS) models from its Hub into a single comparison table. The team explicitly stated there's no "overall best" — and that admission matters more than any specific ranking.

Four objective metrics: WER/CER (word/character error rate — lower means clearer pronunciation; for Chinese, look at CER), RTFx (inverse real-time factor — higher means faster batch generation), TTFA (time to first audio — shorter means smoother conversation, which voice agents care about), and SIM (voice similarity, computed via cosine similarity of WavLM speaker embeddings — higher means closer match to reference audio).

Selection uses a Pareto frontier — a model only displaces another if it's no worse on every metric and strictly better on at least one, sidestepping "weighting averages into a fake champion." Survivors are then filtered by business hard-thresholds (first-segment latency < 250ms, CER < 6%, runs on a single GPU) before going to human blind listening.

The team also drew clear lines: WER only measures intelligibility, SIM only measures identity preservation — neither directly equals "sounds natural" or "conveys emotion."

Industry view

On the positive side, breaking TTS "quality" into four auditable dimensions gives at least a baseline against which to test "our model is best" claims, reducing pure marketing puffery.

But the risks and counter-voices are equally clear:

  • The evaluation script was promised "open source soon" but remained unpublished at press time — no one can reproduce results, so the leaderboard's credibility takes a hit;
  • Low WER ≠ high naturalness — it may just mean clear articulation with mechanical pauses; Chinese additionally requires separate review of CER for numbers, dates, and mixed Chinese-English reads;
  • Short TTFA may come from extremely small chunks enabling frequent scheduling, which actually causes playback stuttering — and interaction metrics like "does the model stop immediately when the user finishes speaking" aren't measured at all;
  • Deployment environments vary hugely: the same model may rank completely differently on H200, consumer GPUs, and CPUs — lab throughput doesn't equal real-world experience.

Our take: this leaderboard's biggest contribution isn't the ranking — it's the evaluation protocol itself. What teams should reuse most isn't the conclusions, but the method: "same hardware, same prompt set, fixed warm-up, report median, segment by language."

Impact on regular people

  • For enterprise IT: This is the first reproducible comparison method for selection. Demand evaluation scripts, scenario-segmented reports, and per-language data from vendors — not just sales talking points.
  • For professionals: People doing phone support, podcasts, or audiobooks can now specify tool requirements more precisely (first-segment latency < 250ms? Chinese CER < 6%?) — no longer swept along by vendor narratives.
  • For the consumer market: the "robotic feel" in smart speakers, car infotainment, and customer service calls won't disappear anytime soon — metric progress doesn't equal experience progress. Naturalness and emotional expression still require human listening.