This week we noticed an interesting development: an independent developer (GitHub: npanj) released an inference engine called Slipstream — a program specifically designed to drive large models to output tokens — and got the 95.5GB open-source model Qwen3.8-Flash-Next running at 41–52 token/s on a 64GB unified-memory Mac, 76% faster than the mainstream solution llama.cpp (the most widely used local inference engine in the open-source community).
The data is what matters. Across 3,086 real programming request tests, speed held steady at 33–44 token/s, and even when context (how much text the model can "read" at once) stretched to 130,000 tokens, throughput didn't drop. On the same hardware at equivalent length, llama.cpp had already slowed significantly.
The project is fully open-source; both the model and the engine can be downloaded directly. It means processing 130,000-token long documents and codebases can now happen locally on a regular-priced Mac.
What this is
Slipstream is, in essence, a C++ inference engine rewritten for Apple Silicon (Apple's self-designed chip). It does two things: first, it shards the oversized model onto SSD and reads-while-it-computes (expert streaming loading); second, it uses speculative decoding (letting a small model guess first, then the large model verifies, reducing computation) to cut compute. The two together deliver speed and stability at once.
Qwen3.8-Flash-Next is the Qwen team's experimental sparse-architecture model — a Mixture-of-Experts (MoE), made of many "sub-experts" with only a fraction activated per token. The 95GB size comes from the total parameter count being large but the per-call activation ratio being low.
Industry view
The bullish camp calls this a milestone for the Apple Silicon AI ecosystem — unified memory architecture (CPU and GPU sharing one memory pool) combined with SSD streaming has, for the first time, given consumer-grade hardware practical speed for "on-device LLMs" (running locally on the device, not depending on the cloud).
But the skeptics are equally clear: a 95GB model being repeatedly read from SSD is fundamentally trading disk I/O throughput for GPU time, putting hidden strain on SSD lifespan and system responsiveness; the project is maintained by a single person, with no SLA, no enterprise support; Qwen3.8-Flash-Next is an experimental model built for a sparse architecture, and its fit for real business workloads remains unverified. The cooler read: this is excellent engineering work, but still two to three years short of enterprise-grade deployment.
Impact on regular people
For individual professionals: from now on, sensitive documents — contracts, internal reports, customer data — can be run through a large model without ever leaving the local machine, eliminating the worry of uploading to the cloud. The catch: you still have to set up the command line, configure permissions, and tune parameters yourself. For now, this remains a perk "for people who know Python."
For enterprise IT: a $2,000–$3,000 Mac workstation can theoretically deliver productivity close to a small cloud model. For industries with strict data-compliance requirements — finance, healthcare, government — this is a cost option worth tracking, but it's not yet ready for production environments.
For the consumer market: the "Mac as AI productivity machine" narrative now has real backing, but it's still far from plug-and-play for everyday buyers. What we're more likely to see in 2026 isn't individuals building their own rigs, but cloud vendors packaging this capability into "one-click monthly" managed services sold to SMBs.