After a year of enterprise AI deployment, this week the developer community trained its sights on an awkward truth: today's RAG (Retrieval-Augmented Generation—having AI search documents first, then answer) still rifles through the entire document library even for "what's 1+1?".
What This Is
RAG has been the de facto standard for enterprise AI over the past two years—feed internal documents to the model so it answers "with evidence," avoiding hallucinations. The problem is that the architecture doesn't discriminate: every question, no matter how trivial, goes through retrieval.
This week, a widely-shared technical post on Juejin (a major Chinese developer community) proposed two patches. The first is a "triage desk": have the model first use structured output (a method forcing AI to produce a fixed-format answer) to decide "is this a common-sense question, or one that needs document lookup?"—simple questions get answered directly, complex ones trigger retrieval, and the path plus data points remain traceable the whole way through. The second is "question decomposition": when a query requires jumping across multiple document passages to resolve, have the model break it into steps first, then retrieve per step, instead of throwing one vague broad question at retrieval all at once.
In essence, this upgrades "mindless retrieval" into "on-demand retrieval plus planned retrieval."
Industry View
Supporters call this the marker of RAG moving from "usable" to "actually useful." What enterprise customers care about most isn't answer accuracy—it's the cost-per-thousand-queries and response latency. Triage can halve redundant retrieval, which translates into real money on enterprise bills.
But cooler heads push back. First, triage itself depends on the model, and the model can also be wrong—misclassifying a "complex question" as "simple" and answering directly is even more dangerous; hallucination (AI confidently making things up) may surface in subtler forms. Second, this architecture is currently only alive in top-tier developer circles; the vast majority of enterprise AI projects are still stuck at "getting basic RAG running"—triage, for them, is a luxury. Third, "triage" is only a patch; RAG's real ceiling—models that can't read long documents, structured data that can't be ingested, multimodal content that's hard to retrieve—remains untouched.
Impact on Regular People
For enterprise IT: If you're evaluating knowledge-base AI tools, ask one more question—"does it have question routing/triage capability?" That's a signal of whether a vendor is going deep, and directly affects long-term operating cost.
For working professionals: When you use an AI assistant to query internal company documents and an answer feels "off," the tool may not be bad—it may have skipped retrieval and just guessed. You can explicitly demand "look up documents first," forcing it down the heavyweight path.
For consumer markets: Smart-customer-service and document-assistant products will become more "tactful"—casual chat won't waste compute, and real questions will trigger deep retrieval. The experience will inch closer to a "colleague who actually knows the field."