An open-source team this week published numbers that are forcing the compute industry to redraw its boundaries: the 177-billion-parameter Qwen3.8-Flash-Next model runs at 9–10 tokens per second on a single 16GB consumer GPU (an RTX 5060 Ti, roughly ¥4,000 in China) paired with 32GB of system RAM and an NVMe SSD. This implies the compute-pricing logic of the past two years may need to be revisited.

What this is

For the past two years, running large models has been almost inseparable from data-center-grade GPUs—an H100 sells for several hundred thousand yuan, and enterprises wanting to deploy a private model first had to clear a million-yuan hardware threshold. As we map it out, the breakthrough this time comes down to SSD streaming.

Qwen3.8-Flash-Next uses a Mixture-of-Experts (MoE) architecture—the model is split into dozens of "small experts," of which only a subset activates per query. Most of the 177B parameters sit idle most of the time, so the team keeps the rarely used "cold" experts on the SSD and reads them into VRAM and system RAM only when needed.

The final allocation: 4.4GB of weights live in the GPU's VRAM, 6–8GB in system RAM, and the remaining ~99GB stays on the SSD for on-demand reads. Each token triggers roughly 480 experts, 75% of which hit memory; a single SSD read pulls about 270MB.

At 9–10 tokens/sec, the speed is already at "readable" levels; prefill hits 49 tokens/sec. For comparison, llama.cpp (the mainstream open-source inference framework) on the same machine manages only 4.9 tok/s—roughly half the speed.

Industry view

We note that supporters read this as another milestone of AI "democratization": compute is no longer the patent of big companies, and mid-sized firms—even individuals—could potentially run hundred-billion-parameter models. But the objections are worth hearing too.

First, the hardware barrier is only superficially low. Support is currently limited to RTX 50-series cards (Blackwell architecture); older GPUs won't run it. The OS has only been tested on Windows 11 and WSL2—Linux server scenarios are uncovered. Decoding is restricted to the simplest greedy method, capping generation quality.

Second, real production environments are not single machines. A Singles' Day e-commerce surge serves hundreds of thousands of users—a home workstation cannot do that. Model training still demands compute clusters; this breakthrough only addresses inference (getting an already-trained model to answer questions).

Third, the toolchain is still fragile. The team itself admits that tool call (letting the model invoke external tools) is not yet stable—there's still a gap before enterprise-grade usability.

Impact on regular people

For enterprise IT: There's no need to scrap existing procurement plans in the short term, but next time GPU cloud services come up for renewal, "Do we really need an H100?" is a question likely to be asked more often—some inference workloads may be enough on consumer-grade hardware.

For working professionals: In another 12–18 months, "running an offline, private AI on your own office PC" may shift from a geek toy into something the average white-collar worker is willing to try—provided GPU and storage get refreshed.

For the consumer market: Cheaper local AI means the partnership landscape between hardware vendors and application companies may reshuffle—but the advantage of cloud giants won't vanish overnight.