This week, a Reddit technical post shared hard data: Apple's M5 Ultra running the Chinese open-source GLM-flash model cut a parameter called prefill step (the chunk size used to pre-process input text before the model responds) from 2048 to 8192, dropping total inference time for a 130,000-token long document from 220 seconds to 185 seconds—a 20% reduction. To us, this isn't a technical footnote. It's a signal that local LLMs have quietly crossed a hidden threshold.
What This Is
The posting engineer was testing Apple's latest M5 Ultra and found that the inference engine (the runtime that actually executes the model) chunks long inputs into smaller pieces for sequential processing—and that chunk size directly dictates efficiency. At the default 2048, a 130,000-token document (roughly 60,000-70,000 Chinese characters) took 220 seconds. Bumped to 8192, it took just 185 seconds. Shorter inputs saw even larger speedups.
The parameter isn't the whole story. Apple's in-house inference framework MLX and the third-party engine omlx have iterated rapidly over the past six months. Models that previously required an A100 or H100—graphics cards priced in the tens of thousands of RMB—are now hitting 700+ tokens/s generation speed on a roughly 40,000 RMB M5 Ultra Mac Studio. That's how many characters the machine reliably spits out per second.
Industry View
Optimists read this as Apple Silicon's formal entry into the AI inference arena—and the beginning of a shift in the local-versus-cloud balance. The number of GitHub open-source projects optimizing for MLX doubled over the past six months. This ecosystem is positioning itself as the second LLM inference substrate after NVIDIA CUDA. The significance, to us, far exceeds a single parameter tweak.
But the cautious voices deserve airtime. A long-time inference-optimization analyst told us over email: prefill optimization yields diminishing returns fast and effectively tops out past 8192; engines like omlx still have unstable compute allocation, with occasional VRAM overflows on long documents. More critically, these performance numbers only matter to developers who read logs and tune parameters. Mainstream office workers opening a Mac and getting out-of-the-box utility remain at least two product cycles away.
Impact on Regular People
- For enterprise IT: Local inference speed is closing in on real-world cloud performance. Over the next year, enterprises can re-run the math on private deployment costs—data-sensitive sectors like finance, healthcare, and legal will move first.
- For working professionals: Technical and creative roles (writing code, producing video, designing) are the first beneficiaries, already shifting some workloads from cloud back to local Macs. General clerical and administrative staff won't feel a difference yet.
- For the consumer market: The M5 Ultra Mac Studio starts at around 40,000 RMB—professional productivity hardware positioning. Average consumers don't need to buy a new computer for "AI localization" right now; this is still a two-to-three-year story.