This week, an unassuming post on r/LocalLLaMA pinned a number in front of us: running Qwen3.8 and several lightweight models on Apple's in-house silicon (Apple Silicon — the M1/M2/M3 series Mac processors) is now up to 3x faster. The poster is a former core contributor to MLX.fast (an open-source project that accelerates AI models on Apple chips), and the new tool is called ishizuki, with code open-sourced on GitHub.
What this is
Simply put: until now, running a language model of Qwen's scale on Mac (the kind of AI that can write articles and summarize text) meant either sluggish speeds unfit for daily use, or routing through cloud APIs (sending your data to companies like OpenAI or Alibaba and paying per call). This optimization makes "fully offline, local execution" look respectable for the first time.
How meaningful that 3x is depends on context: the smaller the model parameters and the older the chip, the more dramatic the speedup. The developer himself emphasizes his best results come on sub-M5 chips — precisely what matters most to the large base of users still running older MacBooks.
Industry view
What's worth acknowledging: in the open-source ecosystem, Chinese open-source models like Qwen are being actively optimized and adapted by overseas developers. China's "open-source for ecosystem" strategy for large models is starting to be validated by the international market with its feet.
The cool-headed counterpoint: the 3x figure is the developer's own claim, not an independent third-party benchmark; the post lives in Reddit's self-promotion section; and the MLX ecosystem (Apple's official open-source ML framework) remains niche compared to NVIDIA CUDA (the industry mainstream). The technical signal is real, but it's far from disrupting the cloud, and traditional enterprise IT won't rewrite its procurement plans over a Reddit post.
Impact on regular people
For enterprise IT: worth tracking, not worth betting on. Production environments still depend on cloud APIs, and local solutions are still catching up on stability, compliance, and observability.
For individual professionals: if you're in a technical role with data that can't leave the company (lawyers, doctors, analysts), running Qwen on a Mac for daily assistance has now moved from "barely usable" to "actually usable." Non-technical roles should still stick with ChatGPT or Claude — less hassle.
For the consumer market: Mac users now have, for the first time, a realistic option for "decent AI without a monthly subscription." But the killer desktop product hasn't arrived — there's no good "Mac AI assistant" wrapping this underlying capability into something a regular person can actually use.