Last Wednesday night, a friend dumped a contract screenshot in my DMs
Last week I was helping a friend review a contract. The clauses ChatGPT generated for him cited three court cases that don't exist. I froze — he'd bought the subscription because some review blogger said "GPT-4 ranks #1 on benchmarks."
What's this got to do with "Benchmarkpocalypse"?
This whole thing traces back to Dan Luu's essay "Benchmarkpocalypse." His point: today's AI leaderboards often crown models that flunk real work — they fabricate legal citations, write broken code, spit out fake data. Why? The benchmark makers got "leaderboard-gamed" — models get specifically trained to ace the tests, then fall apart when actually doing the job.
For us running side hustles or one-person companies, this is the easiest trap to fall into: picking tools based on "which one scores highest," not "who's actually getting real results in their workflow."
I know a freelance translator named Xiaochen. He used to draft business emails with a domestic Chinese model that bragged about "ranked #1 for Chinese." Then a client replied: "Which version of those contract clauses are you citing?" He realized the model had invented three non-existent regulatory codes. Since then he only trusts his friends' real cases — never the leaderboards.
Cost to copy this approach today
Money: $0. You're not buying a tool — you're changing how you pick one.
Time: 1-2 hours. Find 3-5 people in your industry who use AI. Ask them what they've screwed up.
Tech barrier: Can send WeChat messages and make phone calls.
Step one: Open your WeChat groups or Moments, post "Who's been using AI for work lately? Any disasters?" — wait for the real stories to roll in.
Advice for friends at different stages
If you're just starting out (no first client yet): Skip the benchmark leaderboards. Directly ask 3 industry friends what they use and what blew up on them. Spend one afternoon doing this — worth more than a week of reading review articles.
If you have 1-2 stable clients: Build a "my screw-up list" for the AI tools you actually use — what questions did clients ask, where did you give wrong answers. Write it down. Next time you pick a tool, you'll have your own yardstick.
If you're scaling (team of 3+): Build a small group of 5-10 people — clients and peers — dedicated to honest talk about AI tools in real use. Monthly lunch beats any benchmark monitoring subscription.
Last thing: benchmark leaderboards aren't useless, just don't make them your only standard. Skipping them is fine too — first, get your actual work shipped.