What This Is
This week, a hands-on article from Juejin (掘金, China's leading developer community) lays out a sharp idea: rather than hoping AI won't slip up, we should deliberately build tricky questions for it. The author used LangSmith to assemble a 12-question test, stacking 7-day returns, 7-day price protection, 15-day exchanges, and 1-year warranties together — the closer the numbers resemble each other, the deeper they sit in RAG systems' confusion danger zone. A final edge-case question — "free shipping over 99, excluding large cold-chain items" — was designed specifically to test whether AI drops the parenthetical exception clause.
The core action has only two steps: first, build the question bank; then run it through every time the system changes, watching scores move up or down. In software engineering this is called "regression testing" — now ported into AI projects.
Industry View
The pro camp: This is the true mark of an AI project going engineering-grade. In the past everyone compared model parameters and benchmark scores; once you actually ship to production, we realize the evaluation system is where the daily workload lives — typically eating 30%+ of project effort. Frontline consensus: the quality of your question bank sets your AI product's floor.
The anti camp: These 12 questions are all "how many days" or "what time" fact-type queries — easy to score. But 80% of real customer questions are actually ambiguous ("how's your service?"). Pretty scores don't equal user satisfaction. A sharper critique: who decides whether the question bank has 12 or 120 questions? The moment business rules change, the question bank becomes a continuously-maintained burden — new cost, not a free lunch.
Impact on Regular People
For enterprise IT: Don't fixate only on model selection. How the question bank gets built, by whom, and how it's scored — that's the real workload.
For working professionals: We expect roles like "evaluation set engineer" to emerge. You don't need a PhD in algorithms, but you do need meticulousness and business sense.
For consumer markets: Next time an AI customer service gives you an inhumanly standard answer, it's most likely pulling from a question bank — not actually being dumb. That's the industry state, not deliberate coldness.