This week on Reddit's hardware community, a user plans to cram two NVIDIA RTX 5090s (the ceiling of consumer GPUs, with a single card's MSRP at RMB 16,000–20,000) into their home Proxmox server (a home virtualization system) to run open-source AI models, while upgrading the CPU from an Epyc 7402P to a server-grade 75F3. On its own, this isn't big news — but the curve behind it is worth us noting: running local large models on consumer hardware is quietly spilling over from the geek circle into "tech managers curious about AI."
What this is
The RTX 5090 is NVIDIA's consumer flagship released this January, with 32GB of GDDR7 VRAM (AI models run in VRAM — the larger it is, the bigger the model you can load). It's currently the only consumer card capable of running 70B-parameter-class open-source models locally on a desktop (more parameters means a smarter model, but also more VRAM needed). Two cards give you 64GB — equivalent to one enterprise-grade GPU.
The 75F3 in the post is an AMD server CPU, distinguished by its high clock speed (4.0GHz), which gives it an edge over the user's existing 7402P (3.0GHz) for inference response speed — when a model answers your question, CPU clock speed directly affects how long you wait for "the first word to appear."
Combined, this machine's specs are equivalent to moving a small AI inference room into the home.
Industry view
Pro-local-AI voices mainly come from three groups: finance and healthcare practitioners focused on data compliance, developers unwilling to keep paying API fees, and heavy users chasing response speed. The open-source community also welcomes this — stronger hardware demands, in turn, drive progress in open-source model compression and optimization.
But the cautious — even opposing — voices are worth listening to more carefully. Hardware depreciation is a severely underestimated cost: the RTX 5090's product cycle may be only 18 months, after which a new generation doubles compute again and old cards become an electricity bill. Hugging Face's own enterprise deployment whitepaper shows that local AI deployment's cost recovery period typically runs 18–36 months — barely an economic case for individual players. Power and cooling are also hidden big costs — two 5090s at full load draw roughly 1kW, plus the CPU and peripherals, running 24/7 is equivalent to adding a small server room to your home.
Our more realistic judgment: now is not the time to act. The 5090 just launched at a premium, and Reuters reports are already circulating about next-generation consumer cards (5090 Ti or 6000-series). The "wait-and-see crowd" never loses.
Impact on regular people
For enterprise IT departments: consumer GPUs entering local AI deployment is a new variable — this category wasn't in the budget before, and now procurement processes and vendor negotiation strategies need rethinking.
For individual professionals: no need to worry about it for now — your work computer doesn't need this kind of configuration, and subscribing to ChatGPT or a domestic large model membership is more cost-effective.
For the consumer market: the RTX 5090's secondhand resale value will become a fascinating grassroots indicator — whether the AI boom is cooling down, secondhand GPU prices often give the answer before financial reports do.