What This Is

This week, a post in Reddit's r/LocalLLaMA is making the rounds: a user ran Qwen3-27B — a 27-billion-parameter domestic open-source LLM (Alibaba's Qwen series) — on four retired mining RTX 3060 Ti GPUs (8GB VRAM each). Power draw held at 110W per card; throughput hit 120 tokens/sec at a 150k context, dropping to ~70 tokens/sec at the full 262k context.The point isn't "fast" — it's that GPUs others consider obsolete delivered a usable workload. Technically, this leans on Exl3 (a quantized inference engine supporting multi-GPU coordination) and community tools like HyperQwen, stitching four 8GB cards into ~32GB of usable VRAM.

Industry View

The bullish case: open-source models — especially domestic workhorses like Qwen — can now do meaningful inference on consumer hardware. Local AI no longer means "you must buy an H100 (datacenter-grade GPU, tens of thousands to hundreds of thousands of USD)." Second-hand hardware plus open-source software is a real path. Privacy-sensitive companies, indie developers, and teams building POCs (proofs of concept) just got a "starts at a few thousand dollars" option.The bearish case: this is the result of an enthusiast spending weeks tuning — not an out-of-the-box solution. Exl3, HyperQwen, Vllm remain jargon to most people; ordinary enterprise ops teams don't have bandwidth to tinker with quantization (the technique of compressing models to smaller sizes) and parallelism. Power bills, thermals, stability — the miner's common sense — aren't in the post. "Usable" and "a replacement for cloud APIs" are two different things.

Impact on Regular People

For SMB IT: the cost curve for local LLM deployment is flattening — but it's still three to five years from "buy and use." Unless you're in a hard-compliance scenario, don't spin up a project for this; just keep watching.For individual careers: not your problem. Using ready-made products like Kimi, Tongyi, and ChatGPT is still the highest ROI. Reassess when someone wraps the above into a plug-and-play package.For the consumer market: the second-hand GPU market may see a wave of "non-gamer" buyers; refurbished mining cards now have a story. DIY builders and hardware channels are worth watching.