This week, a "monster rig" on Reddit sparked discussion: 6 Bitcoin mining cards running the Qwen LLM, hitting 100k context at 28 tokens/second, with a total power draw of around 115 watts and cooling handled by three fans and a cardboard box. It looks like a geek toy, but it reflects a real trend — mining cards discarded after Bitcoin halving are becoming the most underestimated compute source in the local AI scene.
What this is
The poster Ok-Breadfruit-3523 used 6 BC-250 cards — relics of the 2017-2018 mining boom. These are essentially GPU motherboards custom-built for mining: per-card compute roughly matches a mid-range consumer GPU, but memory bandwidth is weak. Four of the cards run an IQ2_XS extreme quantization of Qwen Next Flash (a technique that compresses the model to a tiny size while sacrificing some accuracy), handling 100k tokens of context at about 28 tokens/second for short text, dropping to 24 tokens/second at 50k tokens. The other 2 cards run a Q4 quantization of Qwen 3.6 35B at 60 tokens/second.
BC-250 cards are dirt cheap on the second-hand market — tens of dollars each, sometimes sold by the pound when mining farms clear inventory. Our read: compute "trickling down" is happening faster than most people expect — not middle-class buyers picking up a 4090, but geeks assembling a 100k-context machine for a few hundred dollars.
Industry view
The bullish camp sees this as a positive signal. Regulars on r/LocalLLaMA noted in the comments that over the past year they've gotten large models running on 3090s, 4080s, even modded P106 mining cards. "Whether Moore's Law is dead, who knows — but compute trickling down is real." In other words, a long-tail market for AI compute is forming: second-hand hardware + quantization + open-source models, all three converging, and the barrier to local AI keeps dropping.
But the skepticism is direct. 115 watts total power draw, 24-28 tokens/second, plus the precision loss from extreme IQ2_XS quantization — fine for chat, but for serious code generation or long-document analysis, it's basically a toy. One AI infrastructure engineer did the math in the comments: at the same electricity cost, a cloud API returns dozens of times more tokens. "Unless you really care about data not leaving your premises, this is just an expensive hobby." Hidden costs also include: reflashed mining card BIOS, aged VRAM, no warranty.
Impact on regular people
For enterprise IT: In the short term, this isn't a cloud replacement. Small teams with sensitive data and tight budgets could evaluate similar mining-card clusters for offline inference on internal knowledge bases — but factor in electricity, ops, and depreciation into TCO. In most cases, cloud still wins.
For working professionals: Almost no direct impact on non-technical white-collar workers. But if your work involves contracts, customer data, or other sensitive content, the "run models locally" option is getting cheaper, and knowing it exists is worth keeping in your back pocket.
For the consumer market: Second-hand GPU prices keep falling. Readers planning a build this year can keep an eye on mining cards, but do your homework: verify VRAM health, reflash to original BIOS, and accept the no-warranty reality.