A technical breakdown lands an uncomfortable number: pure vector retrieval (converting questions into numeric coordinates and matching them to documents by similarity) tops out at a recall ceiling of just 60-70%. Enterprise AI knowledge bases miss roughly 3 of every 10 questions—work routinely blamed on the model being "dumb." But as we'll argue, the real pathology sits one step earlier, in retrieval.
What This Is
Every enterprise AI customer-support bot, knowledge base, and AI-search product follows the same two-step playbook: first find the source material, then let the LLM compose the answer. Step one is called Retrieval; the share of truly relevant documents that get surfaced is called recall.
The judgment we've watched harden across the past six months: relying on vector retrieval alone, recall tops out at 60-70%. The 30-40% that falls through the cracks tends to be the hard stuff—proper nouns, version numbers, product model numbers, internal acronyms. Exactly what enterprises most need to get right.
The fix is Hybrid Search: vector retrieval plus BM25 (a keyword-matching algorithm dating back to the 1970s). Vector retrieval excels at semantic similarity; BM25 excels at exact literal matches. Each returns 50 candidates; the results are merged; a Reranker model re-scores them before handing the top set to the LLM. In our reading of the benchmarks, this lifts recall from 64% to 91%.
Industry View
The bullish camp sees this as an engineering-arbitrage window. The marginal returns from swapping models are diminishing, while the ROI on getting retrieval right is climbing higher. For a mid-sized SaaS company, a 2-4 week rebuild can cut wrong-answer rates by 25-40%—halving inbound customer complaints, by the practitioners we've talked to.
The skeptics counter that this is over-engineering. A Reranker adds GPU cost: re-ranking 50-100 items tacks on 100-500ms of latency, which high-frequency consumer (toC) search users will feel. Tuning the BM25-to-vector weight ratio has to be redone for every business; there's no silver bullet. For products where questions are scattered and call volume is low, the smarter move is to first nail the prompts and document chunking before reaching for Hybrid.
Impact on Regular People
For enterprise IT: over the next year, when you're evaluating AI knowledge base vendors, "how do you measure recall, and is there a Hybrid pipeline" needs to land on the RFP checklist. The era of judging vendors on demo sizzle is over.
For working professionals: when your enterprise AI search gives you a nonsense answer, push back with "was the retrieval step the one that missed it?"—stop blaming the model for being dumb.
For consumer markets: consumer AI search products (Perplexity, Metaso, Nano Search) are heading into another shakeout—the products that get retrieval right will pull ahead in specialist benchmarks.