Last Tuesday afternoon, Ah Wei's thread dropped again

I'm the type who wants to automate everything with AI. Back in March, I had AI follow up on new leads for me. Last Tuesday afternoon, it got to a Guangzhou customer named Ah Wei — after four rounds, it suddenly said "I need more information to continue" and stopped. I stared at the screen for a full minute. I could have kept that conversation going myself, but it couldn't.

You've probably hit this too: AI halfway through starts making excuses — "I suggest you confirm this" or "this step requires your input." I used to blame my prompts. Turns out some of the blame is the model's stamina, not mine.

What is Grok 4.6, and why should we care?

This week xAI dropped Grok 4.6, positioned as a "GPT 5.6-tier agent built for long tasks." The one-line pitch: it won't bail halfway through. Regular models start punting after 8-10 steps — "you should probably handle this manually" — Grok 4.6 claims it can run a complete business workflow without dropping the ball.

From people already using it: Xiaolin runs a maternal content account in Hangzhou. She told me her old model would crash every few days running "collect reviews → categorize → draft reply." She switched to Grok 4.6 to test, and got through 30 reviews without a flip — her words, not an ad. I asked her what tools she'd switched to lately.

Can you try it today?

  • Cost: Grok API has a free tier. Beyond that, roughly 1 RMB handles 50 small tasks — at least 50x cheaper than hiring a part-time support person.
  • Time: Sign-up + my first call took me 40 minutes (I'd never touched an API before). If you're already using ChatGPT-style tools, 10-15 minutes is enough.
  • Tech barrier: You need an API key (the model gives you a remote key), then plug it into Make / Zapier / Dify. Basically teaching the workflow tool to recognize your key.
  • First step: Sign up at console.x.ai, grab your key, then go to Make.com, search "Grok", and wire up your 3 most-asked questions for a test run.

Don't jump in right away — bookmark it, try it when you have time at month-end. Not everyone needs this tool.

Which stage are you at? Here's how I'd choose

Just starting (0-3 customers): Skip it. You're handling 2-3 support messages a day. AI bailing halfway isn't different from you replying yourself. Wait until you're stable at 10+ orders a month.

1-2 steady customers: I'd test it directly. Pick your 3 most-asked questions, let AI handle 50 messages while you review which ones it flubs. Practice in low-risk scenarios first.

Scaling (20+ conversations/day): One bail here costs you a sale. I'd run parallel — keep your current model as primary, route Grok 4.6 as backup. Whoever drops first cedes the lane, so you're not up at midnight when one channel breaks.