Bilibili open-sourced a full translation model suite this week, covering 150 languages with the ability to preserve the original speaker's voice for multilingual dubbing — the moat around localization and translation outsourcing is starting to crack. This is not yet another toy LLM; it's a practical tool that goes straight at this industry's livelihood.

What this is

Bilibili's Index team released Index-Translate. Three parameter scales: 2B, 9B, and 35B-A3B (preview version, 35B total / 3B active parameter MoE architecture), all open-sourced under Apache-2.0 — commercial use and local deployment allowed.

Core capabilities fall into four layers:

  • 150-language text translation, with support for glossaries, writing style, and output format (preserves JSON fields, placeholders, product names)
  • Index-NativeLong: long-document translation, maintaining consistent terminology and names across paragraphs
  • Index-Homura: subtitle/dubbing length control, fitting translations into the timeline
  • Index-Echo: multilingual dubbing using the original speaker's voice (voice cloning)

Translation is no longer just rendering sentences — it's packaging document structure, subtitle timing, and voice cloning together.

Industry view

Optimists argue that localizing a game or producing a multilingual video used to require three separate parties — translators, voice actors, and subtitle teams — and now a single open-source model can collapse it into one step. For SMEs with multilingual needs but limited budgets, this is a cost-cutting weapon.

The risks and controversies are equally clear:

  • Index-Echo's "voice preservation" is fundamentally voice cloning — actors, podcasters, and influencers get their voices cloned into other-language versions, with copyright and publicity rights essentially a legal gray area under Chinese law
  • 35B-A3B demands heavy VRAM; most on-prem users can realistically only run the 9B version, which still lags professional translators on low-resource languages and literary texts
  • The real moat in localization isn't the engine — it's domain glossaries, client review workflows, and field-specific know-how, none of which open-source models can directly replace

Our take: open-source translation models will first eat into standardized, low-margin bulk translation jobs (manuals, product descriptions, short-video subtitles); high-end localization and literary translation are unlikely to be displaced in the short term.

Impact on regular people

  • For enterprise IT: companies with multilingual content needs can deploy on-prem and cut outsourcing budgets, but must still budget for glossary maintenance and review workflows
  • For individual careers: the "language barrier" in translation, foreign trade, and cross-border e-commerce roles is being pushed even lower. Pure translation execution roles will shrink, while people who combine domain expertise with AI tool skills will become more valuable
  • For consumer markets: low-resource-language content on Bilibili and similar platforms will explode, but low-quality machine dubbing and the "thousand voices sounding alike" cloned narration will flood in alongside