What this is
This week on Reddit's LocalLLaMA community, a post by user NickCanCode gained notable traction. He tried running Alibaba's Tongyi Qianwen Qwen3 (27 billion parameters, quantized) on two consumer GPUs (e.g., two RTX 4090s, totaling 48GB of VRAM). His tool of choice was llama.cpp — the most widely used open-source local LLM inference engine. No matter how he adjusted tensor-split (the parameter controlling how model weights are divided across multiple GPUs), at least 1GB of VRAM sat unused. Tuning the weights was like riding a seesaw: nudge the value slightly to the left, and the left card's VRAM spikes while the right card's drops even further.
He suspects the issue lies in MTP (multi-token prediction) tensor-parallel computation being squeezed onto a single GPU. Commenters also noted that VRAM usage for auxiliary components like multimodal projectors is hard to predict precisely, and integer-percentage precision in tensor-split isn't granular enough.
Industry view
The local-deployment community is no stranger to these pitfalls. llama.cpp's multi-GPU scheduling logic has long been characterized as "functional but inelegant," especially when draft models (smaller models used to accelerate inference) and vision modules enter the picture. Some push back, calling it a "niche issue": most local users run on a single card, or simply buy a 48GB single card (like the RTX 6000 Ada) to sidestep splitting.
We think the counterargument deserves equal airtime: open-source community sentiment is easily amplified by technical frustration. One seasoned developer noted that when local deployment requires hours of parameter tuning, the "cost savings" advantage evaporates — cloud APIs have been rapidly improving in both experience and price over the past two years. Our more honest read: local AI is a toy for geeks, not the default option for enterprises.
Impact on regular people
- For enterprise IT: Unless you have hard data-compliance constraints, sticking with cloud APIs remains far more cost-effective than maintaining an in-house local-inference operations team.
- For working professionals: "Running AI on your own computer" sounds cool, but in 2026, the reality is — most laptops can't handle it, and understanding the parameters above is even harder.
- For consumers: NVIDIA, Apple, Ollama, and LM Studio are all working to simplify this, but their efforts only work for entry-level models. Want to run 10B+ parameter models? Brace yourself for pain.