Last Wednesday at 9 AM, I almost lost a deal

Last Wednesday at 9 AM, I was taking my kid to the hospital. A client messaged wanting a video call. I froze for 5 seconds. If you're a solo operator too, you've had that "ambushed by a video meeting" moment plenty, right? I later chatted with Lao Zhou, who runs his own consulting practice. He says he turns down 20+ video invites a year, all due to time conflicts. For us one-person bosses, time gets chopped into pieces, while clients assume "a quick video doesn't cost you 30 minutes."

So what is an AI video twin, exactly?

I just saw something new that made me rethink this. US company Tavus just released a model called Griffin. The gist: build an AI avatar that looks and sounds like you, makes eye contact, listens, and chimes in — basically showing up to your video calls on your behalf. In tests, 48% of people walked away thinking they'd been talking to a real human. That number used to be 2.4%. In other words — within a year, "letting AI take your meeting" might not be sci-fi anymore.

Who's watching this? I know A May, who runs a paid knowledge account on Xiaohongshu. She already uses similar tools to record lessons because "my throat gives out after 4 hours of recording a day." And Lao Qian from a cross-border e-commerce group uses AI avatars for English product intro clips, skipping the cost of hiring foreign presenters. They're not chasing trends — they're one person doing three jobs and need a stand-in.

But I also got stuck — I tried another company's AI avatar before, and the "me" it produced had mismatched lip movements and a wandering gaze. Clients caught it in a heartbeat. So these tools currently fall into two camps: broadcast type (one-way, reads what you script) and conversational type (two-way real-time, listens then responds). Tavus's offering is the latter, currently in closed beta by invite.

What does replicating this cost right now?

Money: Griffin hasn't published pricing yet, but comparable tools run: a 30-second "you" sample clip + $30-100/month (about 200-700 RMB).

Time: recording the sample + tuning is 2-3 hours, no daily babysitting. Once set up, you can generate dozens of videos a month.

Technical bar: zero code, but you need to film a "look sample" — find good lighting and read a script into the camera. I messed this step up — my first recording had a messy background and the AI learned the curtain pattern. Deleted and redid.

First step: to try without spending, search "D-ID" or "HeyGen" — both have free trial versions. Upload a front-facing photo + 30 seconds of audio and see how your "digital self" looks and talks. No card needed, done in 3 minutes.

Three types of people, three uses

If you're just starting out (monthly revenue under 10K RMB): honestly, skipping is fine for now. You still need real humans to build trust with clients — AI avatars right now lean on the "wow factor," and conversion might not beat you showing up on camera yourself. Figure out how to talk to clients first; the new tools can wait.

If you already have 1-2 steady clients: I'd point you toward content compounding — take highlights from one livestream, break them into 10 short clips, let your AI avatar "re-deliver" them, saving you repeat recording time. Skip it for client calls though — risk is too high, getting caught costs you points.

If you're scaling (3-5 person small team): worth seriously looking at for standardized Q&A — pre-sales FAQs, client training videos, give every team member a "twin." But you must add a subtitle at the start: "This is an AI assistant" — once trust breaks, the cost of winning it back is 10x the cost of acquiring a customer.