An indie developer recently discovered in a real-world project: after his team's AI knowledge base went through multiple iterations, nobody could answer whether retrieval had actually gotten better or worse—they could only eyeball a handful of Q&A results to draw conclusions. This points to an uncomfortable reality: the vast majority of enterprise AI knowledge base projects on the market can't even articulate how well they're performing.
What this is
RAG (Retrieval-Augmented Generation—simply put, giving AI a database so it can look things up while answering) is currently the mainstream approach for enterprise knowledge bases. Working on phase two of the Cisu Knowledge Base project, the developer found that after tweaking chunk size, adjusting vector-versus-keyword weights, and swapping embedding models (tools that convert text into numerical vectors AI can understand), evaluation was almost entirely done by "eyeballing a few Q&A results."
He implemented a closed-loop evaluation pipeline: prepare test dataset → run retrieval automatically → calculate four core metrics—Precision@K (how many top results are correct), Recall@K (how many correct answers get retrieved), MRR (how high the correct answer ranks), and nDCG@K (a composite ranking quality score) → log to database → visualize and analyze. Retrieval also uses a hybrid approach—vector recall weighted at 70% and keyword at 30%—balancing semantic understanding with precise matching. He specifically flagged one pitfall: the annotated "correct answer ID" must match the actual chunk ID in the database; otherwise all metrics will read zero, and it's nearly impossible to pinpoint what went wrong.
Industry view
Supporters see this as a landmark step for AI projects moving from "demos that look good" to "engineering deliverables." Without quantification, there's no optimization direction—and no way to prove to the boss that the money was well spent.
But the objections are equally clear. One view: this system is workable in a developer's own project, but at the enterprise level—across departments and vendors—the cost of "manually annotating ground-truth answers" alone is enough to give anyone pause. For a mid-sized enterprise knowledge base with thousands of chunks, a single annotation pass can take dozens of person-days. A more measured line of skepticism: good metrics don't equal satisfied users. We've also heard senior product managers mention that they've seen too many AI projects with "high evaluation scores but users cursing in the trenches"—the problem usually being that the metric definitions themselves are detached from real usage scenarios.
Impact on regular people
For enterprise IT: if your company is procuring or building its own AI knowledge base, the first question you should ask the vendor is "how do you measure retrieval performance, and do you have historical evaluation data?" Those who can't produce numbers most likely don't know themselves.
For individual professionals: the enterprise AI assistants you use daily (the ones embedded in DingTalk, Feishu, and WeCom) likely have never had their retrieval performance systematically evaluated. The odds you'll hit a pitfall are higher than you think.
For the consumer market: in the next 1-2 years, AI knowledge base products will shift from competing on "does it work at all" to competing on "can results be quantified." Vendors that proactively publish evaluation data will be more worth partnering with long-term than those who can only do demos.