This week on Reddit r/LocalLLaMA, a hot post caught our attention: a DIY builder planned to upgrade his local AI server from 4 RTX 3060s to 8, hoping to push Qwen 27B output from 50 tps (tokens per second, roughly equal to characters per second) close to double. The comments, however, were almost unanimously pushing back — past 4 consumer cards, you hit the PCIe bandwidth wall, and throughput may drop instead of rising.

What this is

During multi-GPU inference (multiple GPUs collaborating to run a single large model), GPUs must constantly pass data between each other via the motherboard's PCIe lanes. Consumer GPUs carry large VRAM but limited per-card compute; past 4 cards, communication overhead eats the parallelization gains. Adding cards shifts from "linear speedup" to "fast then slow," and beyond that, just plain slow.

For enterprises trying to cut cloud bills by moving models in-house, "buy more cards = more compute" is a common illusion we want to flag.

Industry view

Supporters argue local AI is cost-controllable at small scale, keeps data on-prem, and has real value for compliance-sensitive sectors like healthcare, legal, and finance.

But we think the more important voice is the opposition: cloud providers use NVLink (NVIDIA's dedicated high-speed GPU-to-GPU interconnect) and InfiniBand (datacenter-grade networking protocol) to sidestep PCIe bottlenecks — something consumer hardware fundamentally cannot replicate. AWS is also still cutting prices on some GPU instances (cloud GPUs rented by the hour) this week. The relative cost advantage of on-prem is narrowing, and "private deployment" risks becoming just an expensive sunk cost.

Impact on regular people

For enterprise IT: Before approving self-built AI cluster budgets, calculate communication overhead first. Don't extrapolate linearly from "card count = compute."

For working professionals: Running models locally remains a viable option for consultants, lawyers, and doctors with privacy needs — but understand the speed ceiling.

For the consumer market: Gaming cards running AI will continue, but for regular users the real impact is on second-hand GPU prices, not "running a ChatGPT at home."