What This Is
This week, an open-source AI voice cloning (TTS, or "text-to-speech") model called Sopro quietly updated to V2: 120 million parameters, runs directly on a regular laptop CPU—the technical barrier is being silently dismantled. First-segment audio latency sits around 300 milliseconds, the code is released under the Apache-2.0 license (free for commercial use), and it currently supports English, European Portuguese, French, and German.
Industry View
Supporters see this as a win for the "lightweight TTS" route—when enterprises build voice customer service or audio content, they can break free from dependence on big-tech cloud APIs, with better control over data and costs. But we notice the other side: open-source plus CPU-runnable means the barrier to voice cloning has been dramatically lowered. Anti-fraud researchers have already warned that AI voice-cloning scams (using a few seconds of sample audio to mimic a relative's voice) will accelerate this year. The author also concedes that extremely high pitches, cartoon voices, and noisy reference audio remain weak spots, and that 120M parameters still leaves a quality gap compared to commercial heavyweights like ElevenLabs.
Impact on Regular People
For enterprise IT: on-premise deployment of voice synthesis becomes viable, and the compliance and cost structure for customer service, audiobook, and podcast scenarios may be rewritten.
For working professionals: content creators and customer-service trainers gain a free tool, but need to watch commercial licensing and compliance boundaries.
For the consumer market: when you receive a call from your "boss" or "family member" asking for a wire transfer, stay alert—the flip side of technology becoming cheap is fraud tools becoming cheap too.