This week, a user on r/LocalLLaMA shared their local AI setup: two RTX 5090 GPUs (consumer flagships, each priced around 14,000 RMB domestically), 64GB DDR5 memory, and an AMD 9950X processor—GPUs alone approach 30,000 RMB. This is no longer a hobbyist toy; it's the starting line for enterprise budgets. They're still debating whether to move this hardware into an enterprise EPYC server. What we care about: as "running AI locally" slides from geek toy to enterprise IT option, which of these three costs—money, ops, or electricity—weighs more?
What This Is
The user wants to run larger open-source models (currently using Flash Next and 27B-parameter-class small models), and is considering moving to an EPYC server—which offers 256GB DDR4 ECC memory and 8-channel bandwidth, with GPUs passed through via PCIe passthrough (using physical GPUs directly inside a virtual machine) to Proxmox (a free enterprise-grade virtualization software, similar to VMware). The question is technical: can the extra memory and bandwidth translate into meaningful inference speed gains?
Industry View
We hear the voices supporting local deployment: data stays on-premises, long-term costs are controllable, no being held hostage by API price hikes. But counterarguments deserve equal airtime:
- Cloud vendors' batch inference pricing continues to drop—Anthropic, Google, and MiniMax are all in a price war.
- Hidden costs of local deployment—dual 5090s drawing ~1kW at full load, driver compatibility, model updates, human ops—are routinely underestimated.
- An AWS engineer replied on a similar Hacker News thread: "When you calculate three-year TCO, cloud is always cheaper—unless your data truly cannot leave the datacenter."
Editorial judgment: only enterprises with hardware budgets in the tens of millions and hard compliance constraints should self-host; for everyone else, paying per token (the basic billing unit for API calls, like "word count") remains the more realistic choice.
Impact on Regular People
- For enterprise IT: the tug-of-war between self-built GPU clusters and cloud APIs isn't over. Only those with budgets in the tens of millions and data-export restrictions should even consider it; for everyone else, pay-as-you-go is more realistic.
- For individual careers: skills like hardware tuning, driver compatibility, and model quantization (compressing models to speed up inference) are hard currency in the local AI scene, but for most roles they're still a plus, not a requirement.
- For the consumer market: this is still far from ordinary consumers. People willing to spend 30,000 RMB running models at home remain a tiny minority on the Chinese internet—this is more of a B2B (enterprise-side) story.