What this is
This week, a build-help post on r/LocalLLaMA drew over a hundred replies: a user tried to fit an RTX 3090 and an RTX 3070 simultaneously into a B450 motherboard to run a local LLM, only to be blocked by the physical gap between two PCIe slots—and a riser-cable workaround clashed with the SATA power cables.
Commenters suggested an NVMe adapter workaround. The post itself is just a personal build question, but it surfaces a truth obscured by industry talking points: consumer-grade motherboards, cases, and power supplies were never designed for multi-GPU LLM (Large Language Model) inference in the first place.
Industry view
The mainstream narrative is "small models on-device + large models in the cloud": Apple Intelligence, Phi-3, and small-parameter Llama models are all pushing in this direction. We've noticed the unspoken subtext—cloud providers and chip companies want consumers to stop touching hardware and hand AI compute over to subscription services.
But the community's counterpoint is clear: to actually run models at 70B parameters or above, or to do local RAG (Retrieval-Augmented Generation—feeding the model your own internal documents), multi-GPU stacking remains the hardcore minority route. The fact that these help threads stay active is itself proof that no mature consumer-grade solution exists outside the cloud.
More alarmingly, hardcore users' real demand has propped up second-hand prices for the 3090 and 4090—LLMs are a hidden driver of the GPU market, a fact sell-side research reports almost never mention.
Impact on regular people
For enterprise IT: deploying LLMs on-prem is not a "plug in a GPU" affair—it's data-center-grade engineering, with racks, cooling, power, and networking all needing rework.
For working professionals: at present, cloud API costs are far lower than building your own. Local deployment suits researchers and tinkerers; regular professionals don't need to worry about it.
For the consumer market: high-end GPU prices won't fall anytime soon—AI training and inference demand is the hidden factor propping up the second-hand GPU market.