A Java tutorial this week surfaces a number we believe every AI project lead should memorize: 0.78 — a similarity threshold for deciding whether the AI "can answer." When a user's question scores below this number against the closest match in the knowledge base, the system should refuse to respond and escalate to a human — not fabricate an answer. This is the most basic, and most overlooked, design pattern in Retrieval-Augmented Generation (RAG) systems.

What This Is

Embeddings convert sentences into arrays of numbers; semantically similar sentences point in similar directions in that numeric space. How closely two vectors align is measured by Cosine Similarity — closer to 1 means more alike; near 0 means unrelated.

Walking through the author's code: using Java 17 to call OpenAI's Embeddings API, they built a small retriever on three FAQ entries. A user asks "how do I exchange a wrongly purchased item," the system converts the query into a vector, compares it against the knowledge base's question vectors, and selects the most similar candidate. If the top score falls below the threshold, the system prompts handoff to a human.

It splits RAG into two stages — "find evidence" and "write the answer" — and emphasizes that "find evidence" carries value on its own, independent of the latest model.

Industry View

The author pushes a counter-intuitive view: with frequent large-model updates lately, many assume "swapping in a newer model will fix hallucinations" (AI confidently making things up). The author argues "retrieval thresholds shouldn't disappear as the name of the generation model changes" — whether GPT-3 or GPT-5, the "what to do when you can't answer" defense line demands serious design.

Counterarguments, however, deserve a cool hearing: the 0.78 threshold is just a teaching starting point, not a universal standard for every language and business. Set it too high and you'll reject questions you should answer; too low and you're back to fabrication. Underneath lies the classic trade-off between recall (fewer misses) and precision (fewer false alarms), which demands iteration — not a one-and-done setup.

Another overlooked risk is engineering cost: the demo recomputes FAQ vectors on every query, but in production these must be precomputed and cached, with the API called only for new queries — otherwise API costs and latency blow up.

Impact on Regular People

For enterprise IT: before deploying AI customer service or a knowledge base, first answer "under what circumstances should the system refuse to respond?" That delivers more ROI than arguing over which vendor's newest model to pick.

For individual professionals: don't be fooled by an AI tool's "confident tone." Whether a sound retrieval foundation sits underneath — and whether it fabricates — is the real test of whether it's worth using.

For the consumer market: when judging an AI customer service, don't just ask "can it answer?" — ask "will it admit it doesn't know?" The latter is the hallmark of a mature product.