Open-source TTS (text-to-speech) model TontaubeV1 beat ElevenLabs Flash v2.5 with a 50.1% win rate in a 400-clip audiobook blind test, and also outperformed paid products like Fish Audio and Cartesia Sonic 3. But don't rush to email procurement: TontaubeV1 currently requires at least 24GB of VRAM to run, and quantized versions are still pending. Worth watching: the open-source camp has breached another gap in the voice AI space dominated by paid oligopolies like ElevenLabs, OpenAI, and Azure—the chips on the enterprise selection table are starting to shift—but the speed and depth of that shift haven't arrived yet.
What this is
TontaubeV1 was released by a two-brother startup team, with 2.9B parameters and open weights, focused on three things: ultra-long text fluent generation (streaming output paired with a rolling context window, theoretically no length cap), zero-shot voice cloning (one minute of reference audio to replicate a voice), and local low-latency inference (0.08× real-time on a single RTX 5090, dropping to 0.02× in batch mode—roughly 50× real-time speed). Technically, it uses four autoregressive models that generate multi-layer codebooks from semantic skeleton to acoustic detail, paired with character-level text tokenization.
Industry view
The optimistic read: open source has breached another gap in commercial fortifications—voice AI is no longer the exclusive preserve of a few paid oligopolies, and enterprise deployment now has a "run-it-yourself" option. But three caveats temper that view. First, the benchmark comes from developer self-evaluated LLM-as-judge (using another large model as referee); Tontaube itself admits large-scale human-ear blind tests haven't been done, and the results haven't been validated by independent platforms like TTS Arena. Second, ElevenLabs' real moat isn't just the model itself but SDK usability, customer support, and compliance audits—open source won the benchmark but may not win the procurement decision. Third, it currently requires at least 24GB of VRAM to run, beyond both consumer hardware and enterprise server rooms; quantized versions are pending. It's more a "open-source hits commercial" signal flare than a near-term replacement.
Impact on regular people
For enterprise IT: Most won't switch in the short term—high-end GPU procurement, compliance, and operations are hidden costs; cloud TTS remains the default option, but the selection negotiation table now has one more chip.
For individual professionals: Self-media creators and audiobook producers will still lean on paid tools like ElevenLabs in the short term, but subscription prices will likely loosen under open-source pressure.
For the consumer market: Local voice synthesis quality on phones and car infotainment systems will improve accordingly, and within a year or two we'll see AI assistants that "sound good even offline."