返回首页

对比阅读

对比阅读:Hugging Face Launches TTS Leaderboard — But Officially Says There's No Winner 与 Hugging Face 上线语音模型评测榜 — 但官方明确说没有第一名

AEN
Hugging FaceOpen TTS LeaderboardText-to-Speech·

Hugging Face Launches TTS Leaderboard — But Officially Says There's No Winner

What this is

On Sept 30, Hugging Face launched the Open TTS Leaderboard, pulling more than 8,000 text-to-speech (TTS) models from its Hub into a single comparison table. The team explicitly stated there's no "overall best" — and that admission matters more than any specific ranking.

Four objective metrics: WER/CER (word/character error rate — lower means clearer pronunciation; for Chinese, look at CER), RTFx (inverse real-time factor — higher means faster batch generation), TTFA (time to first audio — shorter means smoother conversation, which voice agents care about), and SIM (voice similarity, computed via cosine similarity of WavLM speaker embeddings — higher means closer match to reference audio).

Selection uses a Pareto frontier — a model only displaces another if it's no worse on every metric and strictly better on at least one, sidestepping "weighting averages into a fake champion." Survivors are then filtered by business hard-thresholds (first-segment latency < 250ms, CER < 6%, runs on a single GPU) before going to human blind listening.

The team also drew clear lines: WER only measures intelligibility, SIM only measures identity preservation — neither directly equals "sounds natural" or "conveys emotion."

Industry view

On the positive side, breaking TTS "quality" into four auditable dimensions gives at least a baseline against which to test "our model is best" claims, reducing pure marketing puffery.

But the risks and counter-voices are equally clear:

  • The evaluation script was promised "open source soon" but remained unpublished at press time — no one can reproduce results, so the leaderboard's credibility takes a hit;
  • Low WER ≠ high naturalness — it may just mean clear articulation with mechanical pauses; Chinese additionally requires separate review of CER for numbers, dates, and mixed Chinese-English reads;
  • Short TTFA may come from extremely small chunks enabling frequent scheduling, which actually causes playback stuttering — and interaction metrics like "does the model stop immediately when the user finishes speaking" aren't measured at all;
  • Deployment environments vary hugely: the same model may rank completely differently on H200, consumer GPUs, and CPUs — lab throughput doesn't equal real-world experience.

Our take: this leaderboard's biggest contribution isn't the ranking — it's the evaluation protocol itself. What teams should reuse most isn't the conclusions, but the method: "same hardware, same prompt set, fixed warm-up, report median, segment by language."

Impact on regular people

  • For enterprise IT: This is the first reproducible comparison method for selection. Demand evaluation scripts, scenario-segmented reports, and per-language data from vendors — not just sales talking points.
  • For professionals: People doing phone support, podcasts, or audiobooks can now specify tool requirements more precisely (first-segment latency < 250ms? Chinese CER < 6%?) — no longer swept along by vendor narratives.
  • For the consumer market: the "robotic feel" in smart speakers, car infotainment, and customer service calls won't disappear anytime soon — metric progress doesn't equal experience progress. Naturalness and emotional expression still require human listening.
来源: juejin.cn
BZH
Hugging FaceOpen TTS Leaderboard语音合成·

Hugging Face 上线语音模型评测榜 — 但官方明确说没有第一名

这是什么

9 月 30 日 Hugging Face 上线 Open TTS Leaderboard,把自家 Hub 上 8000 多个语音合成(TTS,即 Text-to-Speech)模型拉进同一张表横向比较,官方明确表态:没有「全面最强」,这件事比具体排名更重要。

四个客观指标:WER/CER(词/字错误率,越低发音越准,中文看 CER)、RTFx(逆实时因子,越高批量生成越快)、TTFA(首段音频时间,越短对话越流畅,语音 Agent 在意)、SIM(声纹相似度,用 WavLM 说话人嵌入的余弦相似度计算,越高越像参考音)。

筛选用 Pareto 前沿——一个模型只有在所有指标上都不比对手差、且至少一项更好,才能淘汰别人,避免「加权平均出一个伪冠军」。留下的按业务硬门槛(首段延迟 < 250ms、CER < 6%、单卡可跑)过滤,再交给人工盲听。

官方也主动划了边界:WER 只测可懂度、SIM 只测身份保持,都不直接等于「听起来自然」「情绪到位」。

行业怎么看

正面看,把 TTS 的「好」拆成可审查的四个维度,至少让「我家的模型最好」有了一个对照基线,避免纯营销自吹。

但风险和反对声音同样清晰:

  • 评测脚本说「很快开源」到发稿仍未发布,现在没人能复现,榜单可信度打了折扣;
  • WER 低 ≠ 自然度高,可能只是发音清楚但停顿机械;中文还要单独看数字、日期、中英混读的 CER;
  • TTFA 短可能靠极小分片换来频繁调度,结果反而出现播放卡顿,而「用户说完模型是否立即停」这种交互指标榜单根本不测;
  • 部署环境差异巨大,同一模型在 H200、消费级 GPU、CPU 上的排序可能完全不同,实验室吞吐不代表真实体验。

我们的判断:这份榜单最大的贡献不是排名,而是评测协议本身。团队最该复用的不是结论,而是「同硬件、同提示集、固定预热、报告中位数、按语言分别看」的方法。

对普通人的影响

  • 对企业 IT:选型第一次有了可复制的对照方法,应该向供应商要评测脚本、分场景报告和分语言数据,而不是只听销售口径。
  • 对个人职场:做电话客服、播客、有声书的人,可以更精确地对工具提要求(首段延迟 < 250ms?中文 CER < 6%?),不再被厂商话术带跑。
  • 对消费市场:智能音箱、车机、客服电话的「机械感」短期不会消失——指标进步不等于体验进步,自然度和情绪表现目前仍要靠人听。
来源: juejin.cn