Back to home

Compare

Comparing: Google Puts a Face on Gemini: Real-Time Voice + Avatar, but Hurdles Remain & Google 给 Gemini 加了张脸:实时语音加虚拟形象,但真正落地还得过几道关

AEN
Google DeepMindGeminiLive Avatar·

Google Puts a Face on Gemini: Real-Time Voice + Avatar, but Hurdles Remain

Google DeepMind launched Gemini 3.8 Live this week, merging real-time voice conversation and virtual avatars into a single model. To be clear, this isn't new territory — OpenAI and Alibaba both demoed similar capabilities in 2024, but Google has productized it more thoroughly this time: latency is officially under 1 second, and developers can run it directly via API (a standardized channel letting external programs call in).

What this is

In short: the AI doesn't just listen and respond — it simultaneously generates a virtual face that makes expressions and adjusts tone to converse with you. You can interrupt it anytime, like a real video chat with a person.

The key isn't doing voice or doing avatars separately — it's having both capabilities synchronized within a single model. Lip-sync, expression, and timing are aligned at the foundation layer, so developers don't have to stitch it together themselves. That's the biggest gap versus past "voice model + third-party digital human" bolted-together solutions.

Industry view

The bullish take: this is a critical piece of the "final form of AI assistants" puzzle. Next-gen customer service, online education, and companion apps will all be rebuilt on top of it. In Google's demos, the AI avatar shifts expressions and adjusts cadence based on context — a major leap in human-likeness over voice-only.

But we've heard plenty of sober voices too. One digital-human startup executive told us privately: "There's a hundred-mile gap between the demo and production." The biggest cost in avatar products isn't the model — it's rendering and bandwidth. The compute cost of one parallel conversation can run dozens of times higher than a text-only chat. And the "uncanny valley" — the instinctive discomfort people feel toward something almost-but-not-quite human — remains unsolved. Many users experience instinctive aversion to an AI face.

There's also a regulatory variable: real-time audio-video generation is bumping up against the boundaries of "deepfake" content (AI-synthesized fake video of real people). The EU AI Act already covers this territory, and Chinese guidance is on the way.

Impact on regular people

  • For enterprise IT: SaaS (subscription software services) in customer service and telesales will integrate first. Mid-sized companies won't need to build in-house digital-human teams anymore — they can just plug into the API.
  • For individual careers: Sales reps, trainers, and livestreamers — jobs built on personal delivery — won't be replaced in the short term, but the tools are shifting. In the future, you may be the one directing an AI avatar to record your course.
  • For consumer markets: Companion and virtual-idol apps will see a wave of experience upgrades, but how many people will actually pay to "face an AI avatar every day" remains an open question.
BZH
Google DeepMindGeminiLive Avatar·

Google 给 Gemini 加了张脸:实时语音加虚拟形象,但真正落地还得过几道关

Google DeepMind 这周推出 Gemini 3.8 Live,把「实时语音对话」和「虚拟形象」合进同一个模型里。说实话这不是新方向——OpenAI 和阿里 2024 年都演示过类似能力,但 Google 这次把它工程化得更彻底:官方称延迟压到 1 秒以内,开发者可以直接接 API(应用编程接口,即让外部程序调用的标准化通道)跑起来。

这是什么

简单说:AI 不只能听你说话并回答,还能同时生成一张虚拟脸,对着做表情、调整语气跟你一起交流。你可以随时打断它,像跟真人视频聊天一样。

关键不在单独做语音或单独做虚拟人,而是两个能力在同一个模型里同步——嘴型、表情、时机的对齐在底层就解决了,开发者不用再自己拼一遍。这是它和过去「语音模型 + 第三方数字人」拼接方案最大的区别。

行业怎么看

支持方的判断:这是「AI 助手最终形态」的关键拼图,下一代客服、在线教育、陪伴类应用都会基于此重构。Google 演示里 AI 形象能根据语境切换表情、调整节奏,比纯语音拟人了一大截。

但我们听到了不少冷静声音。一位数字人创业公司高管私下跟我们说:「Demo 和上线之间隔着十万八千里。」Avatar 类产品最大的成本不在模型,在渲染和带宽——一个并发会话的算力开销,可以是纯文本对话的几十倍。另外,「恐怖谷效应」(人面对像人又不完全像人的东西会本能不适)至今没解决,很多用户看到 AI 脸会出现本能抗拒。

还有一层监管变量:实时音视频生成撞上「深度伪造」(用 AI 合成虚假人像视频)的边界,欧盟 AI 法案已管到这一块,国内相关指引也在路上。

对普通人的影响

  • 对企业 IT:客服、电销行业的 SaaS(软件订阅服务)会率先整合,中小公司不用再自建数字人团队,接 API 就能用。
  • 对个人职场:销售、培训师、主播这类靠表达吃饭的工作短期不会被替代,但工具正在变——未来你可能是指挥一个 AI 形象去录课的人。
  • 对消费市场:陪伴类、虚拟偶像类 App 会有一波体验升级,但「天天对着一张 AI 脸」究竟有多少人真的买单,仍是开放问题。