What this is

This unfolded on Reddit's LocalLLaMA board. Community team Yamz Labs built a local inference engine called Kyojin on top of ExLlamaV3 (a mainstream open-source inference engine) and AMD ROCm (GPU compute platform), purpose-built for AMD's Strix Halo platform (Ryzen AI Max+ 395, with 128GB unified memory).

They quantized and compressed two 30-billion-parameter MoE (Mixture of Experts — a large model where only a subset of parameters activates per inference) models to fit this mini PC: GLM-5.3-Flash at roughly 99.7GB, hitting 580 tok/s prefill and 26-30 tok/s generation; MiMo-V2.6-Flash at roughly 105GB, hitting 650 tok/s prefill and up to 44 tok/s generation (in code scenarios).

The point isn't that it "runs" — it's that it "runs well." 44 tok/s is roughly half of normal human reading-aloud speed, nowhere near the hundreds-per-second that cloud APIs deliver, but already usable.

Industry view

Supporters are calling this a milestone for "local AI" — enterprises no longer have to ship all data to the cloud; a single mini PC can run a 300B-class model for internal document processing or coding assistance, with stronger privacy and tighter cost control. The dirty work — quantization, ROCm adaptation, inference engines — that only big labs used to bother with is now being picked up by the open-source community.

But we have three reservations. First, KLD divergence (a metric measuring how much the quantized version's outputs deviate from the original) shows visible drift from the official FP8 (8-bit floating point) versions — GLM's top-1 token consistency is only 89.3%. Second, the Strix Halo hardware ecosystem is still small; an average user still needs serious technical chops to buy the machine, install drivers, and get a pipeline running. Third, 44 tok/s is still a long way from "smooth conversation," and domestic cloud APIs remain the more cost-effective choice; the sustainability of this approach is also uncertain — Yamz Labs is a community team whose iteration depends on individual effort.

Impact on regular people

For enterprise IT: Scenarios involving sensitive data that don't want to ship corpora to public clouds now have an additional local-deployment route. But it's far from "out of the box" — at minimum, you need an engineer who understands ROCm and quantization.

For individual professionals: For daily email writing, translation, or business Q&A, going straight to Kimi, Tongyi, or Ernie is still faster and cheaper; local deployment matters more for programmers and AI hobbyists.

For the consumer market: AMD and OEMs will most likely double down on the "AI mini PC" category through 2026; if prices drop below 8,000 RMB, ordinary users might actually pay up for "a PC that can run large models."