Here's what we're watching: a developer hand-tuned the inference engine kernel (the lowest-level execution rules) for the 40-core M5 Max, pushing local large-model generation on Mac up another 20%. The number doesn't grab headlines—but it surfaces a cost Apple has been burying: every silicon generation demands manual rework to stay competitive in AI.

What this is

The engine is Splash—a local inference program that runs large language models on your machine without cloud round-trips. Its default parameters were originally tuned on 16- and 20-core M5 chips, plus the 32-core M4 Max. Drop that workload onto the 40-core M5 Max and those defaults are no longer optimal.

The developer applied a "split-K" layout—splitting each computation row into four parallel chunks, then merging the results—and specifically optimized the text-generation step (the industry calls this decode, the process of spitting out answers token by token). That's where the +20% came from.

The tool is open-source (GitHub: incoai/splash), but the bar isn't low: macOS 26 + Xcode 26 required, must build from source, and automated tests still need human verification. It's still firmly a geek-tier tool.

Industry view

Bull case we buy: Apple Silicon has structural advantages in AI inference. Unified memory lets CPU and GPU share large VRAM pools; the performance-per-watt ratio is strong; this is one of the few consumer hardware tiers that can run 70-billion-parameter models on a laptop or workstation. The market position is scarce.

But the risks are equally clear:

  • Not a free lunch: That +20% was bought with per-generation hand-tuning. Regular users can't run this play.
  • Speed dividend evaporates fast: Models themselves iterate quickly. Once a new generation lands, Splash's tuning gains may flatten within months.
  • Ecosystem fragmented: Apple supports only its own Metal API (the analog to NVIDIA's CUDA). Developers can't write one codebase to cover every platform—workload effectively doubles.

Our read: hardware wins at the starting line. Software ecosystem is dragging behind.

Impact on regular people

For enterprise IT: Procurement checklists gain a new line—can it run 10B+ parameter models locally? It's becoming a buying criterion, especially for sensitive-data workloads where cloud isn't an option.

For working professionals: An M5 Max Mac around $4,200 (≈¥30,000) running local large models is moving from "geek toy" to "viable option"—but it still asks you to compile code. Plug-and-play remains out of reach.

For the consumer market: The "AI PC" narrative for Mac gets fresh ammunition. But the near-term audience is still a developer niche. If Apple wants this road to actually open up, we think the toolchain needs to drop another notch—the way it once courted iPhone app developers.