What This Is
Developer roandejader has local RAG running on 1.2GB of VRAM—and his solution is: throw out the vector database entirely. The open-source project Hillock writes structured facts directly into SQLite, then uses 10,000-dimensional "hypervector gates" to match and verify queries. Irrelevant questions get blocked before the LLM ever generates a response, stopping hallucinations at the source.
During ingestion, a lightweight bi-encoder extracts facts like names, specs, and dates—bypassing the generative model entirely—taking about 5 seconds. At query time, the CPU runs gate checks via popcount instructions in just 0.01 milliseconds. The full stack plus an 8B model stays under 1.2GB of VRAM, saving several times the footprint of a typical Chroma + 8B setup. It's OpenAI API-compatible and plugs into Open WebUI, AnythingLLM, and Obsidian.
Industry View
The significance isn't "another RAG tool"—it's that Hillock challenges a default assumption: that local knowledge bases require vector databases. It hits a real pain point. When running an 8B model locally, Chroma often consumes VRAM comparable to the model itself. And traditional cosine similarity is genuinely weak at hard-negative rejection—treating questions that aren't in the document as relevant, which leads the model to confidently fabricate answers.
But the counterarguments are clear. First, this is a solo project with no scale validation; enterprise-grade RAG needs Pinecone/Weaviate-style multi-tenancy and permission management, and won't be displaced short-term. Second, the author openly admits the system prioritizes precision over recall—summarization and open-ended Q&A get rejected outright. Third, vector databases themselves are evolving: hybrid retrieval and reranking are already closing the gap.
Our judgment: Hillock won't reshape the mainstream architecture, but it reminds us that the AI infrastructure layer is far from settled—and that garage-level innovation can still challenge defaults we've taken for granted.
Impact on Regular People
For enterprise IT: Short-term, ignore it. Most corporate knowledge bases still rely on mature vector databases.
For individual professionals: Tech enthusiasts who've wanted to give their local AI a knowledge base but been put off by VRAM requirements can now try it—just pip install hillock.
For the consumer market: If lightweight approaches mature, cheaper local AI knowledge products could emerge—but we're still at least one to two years out.