Scene hook

When I was still revising a client proposal at one in the morning, what scared me most wasn’t slowness. It was switching to a “powerful” new model and ending up with more chaos the longer I used it. I’ve been stuck there too: I’d see a model shoot up the rankings and immediately want to try it, only to realize the next day that what really delayed the work was unstable output.

What this tool/method is + who’s already using it

I took this as a very simple reminder: don’t just look at model test scores; look at whether it actually feels usable in real work. On the evening of July 16, in a coffee shop in Binjiang, Hangzhou, Aning, who sells knowledge products, was replying to student messages while showing me results from a model she had just tested. It was fast at summarizing long text, but the moment she used it to revise sales copy, it started to drift. She said on the spot, “It may look good, but that doesn’t mean it’s useful.” That really landed with me, because I’ve made that mistake before too.

What it costs to replicate today

If we want to try this ourselves, the replication cost is actually low: money, 0-30 RMB; time, 20 minutes; technical barrier, basically just being able to open a webpage and paste in a piece of our own real material; first step, click “Start Chat” or “New Chat” on the model page. I’d suggest not testing it first with “write a poem” or brainteasers. Use something we actually need to deliver today instead, like a client reply, a pricing explanation, or a course outline. That way we can tell right away whether it actually saves us time. And if we don’t want to try it right now, that’s fine too—this isn’t a tool everyone needs.

Advice by stage: just starting / 1-2 clients / scaling up

If we’re just starting out, I’d treat it as a “drafting assistant” first—use it to organize ideas and write a first version of copy, without rushing to chase the newest model. If we already have 1-2 clients, I’d focus on testing how consistently it handles proposal revisions, follow-up messages, and meeting-note summaries, because those directly affect conversions and delivery. If we’re scaling up, I’d care more about whether everyone on the team can use the same prompts and produce stable output, rather than who got to the latest model name first. For a small team, avoiding one extra round of switching costs is often worth more than gaining a few benchmark points.