This week a developer ran an 85GB DeepSeek-V4-Flash model on a 12GB RTX 3060 at roughly 3 tokens/s. On the surface it's a tinkerer's experiment; our read is that local-LLM's cost wall is loosening, and the default assumption that "AI must run in the cloud" is starting to wobble.
What this is
A typical gaming PC has only 12GB of VRAM, which traditionally can't hold an LLM sized in the tens of GB. Developer Ivan Adriazola's approach: use an NVMe SSD as a fourth memory tier. Because MoE (Mixture of Experts) only activates a fraction of parameters per inference, SSD reads stay manageable. Combined with madvise(WILLNEED) — a syscall that tells the OS "prefetch this data, it's about to be used" — Linux pulls multiple expert layers into memory in parallel. Result: on identical hardware, the model's expert-loading speed is roughly 4x faster than running llama.cpp (the mainstream open-source inference engine) naively.
Industry view
Optimists call this the "meant-for-it" optimization path for MoE — frontier labs are busy stacking H100 clusters, but the local community is going a different way: making consumer hardware run these models.
Rather than cheerleading, we'd flag three risks. First, 3 tokens/s is still slow for chat — a full AI reply takes over ten seconds. Second, cold start (first model load) takes 102 seconds, and that's the wait every time you switch conversations. This is an engineering experiment, not a product. Third: single machine, single model, single developer's self-test, not yet community-verified — the author himself admits a heavy dose of "vibecoding" (code-by-vibes), so the conclusion shouldn't be generalized.
Impact on regular people
For enterprise IT: if the approach replicates, the cost of running LLMs locally on sensitive data drops significantly, giving data-compliance scenarios another option.
For individual professionals: still a geek toy today, not a productivity tool — but the signal is worth watching. The cost curve for local LLMs has started to bend.
For the consumer market: if "local AI" becomes a trend, GPUs and high-capacity NVMe SSDs may once again become hot hardware.