An overseas user has run Alibaba's Tongyi Qianwen team's Qwen3.8-Flash-Next model on an RTX 5090 (a consumer-grade GPU retailing for roughly 16,000 RMB), achieving a decoding speed of 40 tokens/sec while consuming only 27GB of VRAM. What we find worth paying attention to: this number means a desktop PC in the 10,000-RMB range can now potentially run a hundred-billion-parameter large model — local AI is shifting from a geek toy toward a practical tool.
What this is
This week, a user on the r/LocalLLaMA forum ran the Qwen3.8-Flash-Next model using a single RTX 5090 paired with an ordinary NVMe SSD. Decoding speed: 40 tokens/sec (a token is the smallest unit of text a model processes, roughly corresponding to one Chinese character), with only 27GB of VRAM used.
The model adopts a "hybrid architecture" — part of its parameters stream from the SSD, rather than requiring everything to be loaded into VRAM (the GPU's onboard compute memory). Large models with hundred-billion parameters, which previously required stacking H100 servers, may now run on a single high-end desktop.
Industry view
Supporters see this as a significant counterpunch from the open-source camp against closed-source cloud offerings: local deployment means data never leaves the premises, monthly fees become electricity bills, and long-term costs are controllable. For compliance-bound industries like finance, healthcare, and law, the appeal is especially strong.
But dissenting voices deserve attention too. A developer who has long tracked local inference points out: at 40 tokens/sec generation speed, generating a thousand characters means waiting half a minute — the experience is far inferior to the streaming responses of cloud-based GPT-4o or Claude. Meanwhile, the total cost of an RTX 5090 plus 64GB of RAM exceeds 20,000 RMB, making the payback period not short. More critically, the model itself is still a snapshot from a few weeks ago, and as cloud models iterate weekly, local deployments must chase that capability gap on their own.
Impact on regular people
- For enterprise IT: SMEs with data sensitivity or compliance constraints now have an alternative that doesn't rely on overseas APIs (application programming interfaces), but operational complexity remains a real issue.
- For individual careers: the "toy threshold" for tech enthusiasts is falling, but turning it into a productivity tool still requires engineering capability. For ordinary white-collar workers, the cloud is still the easier choice for now.
- For the consumer market: high-end GPUs and large-capacity SSDs remain beneficiaries in the short term, but if local AI truly takes off, the pricing power of cloud subscription services will be eroded.