What This Is

This week we noticed a small story with a big implication. An open-source model called Gepard posted a 68.7 ms time-to-first-byte latency on the third-party benchmarking platform Coval—faster than all 24 paid APIs on the leaderboard. It runs on a single consumer-grade RTX 4090 GPU (retail price around ¥15,000 / ~$2,000), and the model is released under Apache 2.0 (commercial use, modification, and redistribution allowed). In other words, voice synthesis that used to require paying ElevenLabs or OpenAI by the second may now run on a single graphics card.

Industry View

Supporters frame this as the "Llama moment for voice"—open-source LLMs caught up to closed-source in 2023, and 2025 is voice's turn. Open-source pulls down the cost floor for audiobooks, short-video dubbing, and IVR (interactive voice response) menu systems, letting small teams build voice products that previously only large companies could afford.

But the counterarguments are equally clear and worth flagging. First, low latency doesn't equal quality—the real tests for TTS are voice clone fidelity, emotional range, and long-form stability; 68.7 ms only answers the "is it fast" question. Second, the benchmark was run by the model team against their own API; Coval didn't independently verify it, leaving room for parameter tuning to game the leaderboard. Third, enterprise TTS procurement also weighs compliance, branded voices, private deployment capability, and reference cases—gaps open-source models won't close in the short term.

Impact on Regular People

For enterprise IT: Voice module costs for customer service bots and smart hardware will keep falling—procurement budgets for the year are open for renegotiation.

For individual professionals: Short-video, audio content, and podcast creators will find the tools cheaper—but the field will also get more crowded. Knowing how to use AI is no longer a differentiator.

For the consumer market: We'll see more AI-dubbed audiobooks and short videos, while human voice actors become rarer and more valuable. The question "is this a real person or AI?" will keep getting harder to answer.