Scene hook

Last night at 10 p.m., I was staring at a “No. 1” tool and almost paid the tuition again.

I know that feeling too well: the client needs results tomorrow, and the moment we see a leaderboard, an award, or a hot post, it’s easy to think, “I can’t really go wrong buying this one.” But I’ve gotten stuck here too. Something can look amazing on a ranking and still fall apart once it’s dropped into real business work.

What this method is + who’s already using it

What snapped me awake this time wasn’t some new tool, but an old problem: some AI systems are great at “scoring well,” but not necessarily great at “doing the job.” Even in high-attention competitions like Kaggle, people have been discussing how AI solutions with pretty thin substance can still win big awards.

Last Tuesday at 3 p.m., I was in a cafe on Wantang Road in Hangzhou, testing copywriting tools with A Ning, who works on lead generation for education businesses. She kept clicking “generate,” then sending the outputs to two longtime clients for a quick look. The tool with the flashiest leaderboard result was the one the clients said felt “machine-written.” What we kept in the end was the option with the lower score, because it sounded more like an actual person. I’ve messed this up too: I used to trust public rankings too much and skipped the step of getting real humans to review a first round.

What it costs to replicate today

To copy this “anti-crash check,” the cost is about RMB 0-99, the time is 20-40 minutes, and the technical barrier is very low: if we can copy and paste and open a spreadsheet, that’s enough. No need to understand the jargon.

The first step is simple: open “New chat” or “Start generating” in the AI tool you’re currently testing. Then test it with 3 tasks you would actually send to a client, and ask 2 real people to review the results: does it look like something you’d really deliver, can it be used as-is, and where does it break the illusion? Not everyone needs to do this right now; if you still have only a few clients, it’s fine not to run the test yet. Just remember one thing first: don’t look only at the leaderboard.

Advice by stage

If I were just getting started, I’d use the free plan to test 3 real tasks first and not rush into an annual subscription.

If I already had 1-2 clients, I’d put “can the client use this directly?” ahead of “how new is the model?” and do a quick human review sample first.

If I were scaling up, I’d keep a shared sheet and track each tool’s pass rate, rework rate, and client feedback. We’re not paying for hype. We’re paying for stable delivery.