This week on r/LocalLLaMA, a developer shared data: he ran Qwen-Flash-Next (q4 quantized version) on his 64GB M5 Mac mini, hitting 17.5 token/s decoding speed — but the GPU sat idle 27% of the time waiting for the SSD to read in the required expert modules.
Behind that number is a fact the industry keeps sidestepping: the hardware bar for running large models locally is far higher than vendors advertise.
What This Is
Qwen-Flash-Next is a MoE (Mixture of Experts) model. Think of it as a company with dozens of specialists on staff — only a few are called for any given question, so while total parameters are huge, every inference doesn't need to load them all.
But even with MoE, the full model still exceeds 64GB. This developer's approach: keep the most-used experts in memory (75% hit rate), stream the rest from SSD — essentially treating the drive as extended memory. The result: SSD is far slower than RAM, and the GPU frequently sits waiting for data.
Industry View
The community is treating it as a milestone: 390 token/s prefill speed and 72% prefetch accuracy are already pushing the limits of consumer hardware — software optimization is forcibly turning the impossible into possible.
But the counter-arguments are equally clear-eyed. Cloud vendors will point out that this approach means unpredictable per-inference latency and poor stability — unusable for real workloads. One AI infrastructure engineer told us privately: running it and running it usefully are two different things; no enterprise will make customers wait on SSD loads. And worth flagging: open-source models iterate so fast that today's optimizations may be obsolete next month.
Impact on Regular People
For enterprise IT: Companies planning local LLM deployment need to re-budget. A 64GB Mac mini, at over 20,000 RMB, is barely adequate — and still far from production-grade.
For individual professionals: Running AI tools on a Mac is viable, but don't expect it to replace the cloud — speed, stability, and cost aren't in the same league.
For the consumer market: The "AI PC" concept vendors are pushing still needs another year or two of hardware iteration before it can actually host top-tier open-source models.