Last week a client asked why my tone had changed
Last week I had AI send follow-up emails automatically. The next day, a longtime client asked me, "why does the tone sound different?" That was my wake-up call — I'm letting AI do so much every day, who's catching the mistakes?
Guidelight just graded the five big AI labs
Guidelight is a brand-new non-profit, and they recently did something interesting: they pulled Anthropic, OpenAI, Google, xAI, and Meta into one room and scored them across six dimensions — monitoring, interception, traceability, permission management, audit logs, and incident response.
The results were telling: everyone's doing well on monitoring, but when it comes to actually stopping AI from doing something risky, there's not much evidence any of them can do it well yet.
I used to assume the big AI labs all had safety nets built in. After reading this report, I realized it's not that solid. A friend of mine, Xiaolin, runs a cross-border e-commerce shop. Last month he had AI handle PayPal refunds automatically — it confused "dispute handling" with "direct refund" and burned $2,000 overnight. When AI takes these agentic actions and gets it wrong, who's your safety net? That's exactly what Guidelight is pushing the labs to answer.
What you can do today: audit what you're letting AI do
- Money: Free. The Guidelight report is publicly available.
- Time: 30 min reading the report + 1 hour auditing what AI does for you.
- Technical barrier: No coding needed, just list-making.
- First step: Search "Guidelight AI control grades" for the full report.
Three concrete moves:
- Make a list of everything you currently let AI do (write copy, reply to emails, read contracts, build spreadsheets).
- Next to each one, write one sentence: "If AI gets this wrong, will I know immediately?"
- For the high-risk ones, add a human spot-check.
My own dumb-but-works method: every Friday afternoon, I manually review 10% of what AI sent out that week. Takes 20 minutes. Has saved me twice — once was a name mix-up, once was a misplaced decimal in a contract amount.
Advice by stage
If you're just starting out (no clients yet / not using AI for work): don't panic. You can skip this for now. If you plan to use AI, start with low-risk stuff like writing copy, making images, organizing notes — don't let AI touch client data or money on day one.
If you have 1-2 clients (already using AI for daily tasks): here's one thing I'd do — build a list of "AI is touching my client data" items, and for each, write "how would I notice if it broke." Guidelight says the labs all do monitoring, but real-world interception isn't there yet — so your own safety net matters more than theirs.
If you're scaling up (team of 3+ / stable monthly income): I'd seriously look at which AI vendor you're picking. Anthropic and OpenAI ranked top two in the report, but they themselves admit their interception gaps. If your business has AI moving money, contracts, or client privacy, add a layer of "human confirmation on AI output" — don't let it run end-to-end on its own.