What this is

We noticed a Reddit user who got a 30-billion-parameter (30B) open-source large language model running on two old devices totaling under 20GB of memory (one mini PC, one Mac mini), hitting inference speeds of 30+ tokens per second (tokens are the smallest unit of text a model processes). This isn't a performance breakthrough — it's a directional signal: as RAM prices are pushed up by data center demand, has "cobbling together old machines to run AI" become a viable second path? The user also split their RAG (Retrieval-Augmented Generation, letting AI retrieve external documents to answer questions) knowledge base across both machines' memory without writing to disk, and designed automatic failover logic for disconnections.

Industry view

The local AI community largely agrees that "pooling memory across a LAN" actually works, but flagged three problems: first, latency collapses over wide-area networks, so this currently only suits small teams; second, splitting documents across multiple machines blurs security, compliance, and privacy boundaries — finance and healthcare scenarios are unlikely to touch it; third, maintenance costs are underestimated — one person managing two devices is fine, but enterprise IT handling two hundred is a different story. We also note that cloud vendors' token prices have dropped more than 10x over the past two years, so this "cost-savings route" currently fits better for scenarios where data must stay on-premises, not for pure cost-cutting.

Impact on regular people

For enterprise IT: repurposing old devices now has a new rationale, but we recommend piloting only in data-sensitive, small-scale departments — don't benchmark it directly against cloud service reliability commitments. For individual professionals: lawyers, doctors, consultants and others handling confidential documents may soon process files with local AI without worrying about client data leaving the company network. For the consumer market: secondhand prices for mini PCs and older Mac minis may get a floor from this use case — "can run AI" is becoming a new selling point.