We saw an interesting number: a developer spent a few days using Anthropic's Claude (currently one of the strongest AI coding assistants) to rewrite the Splash open-source inference engine, specifically optimizing it for Apple's M5 Max chip. The result — up to 50% faster concurrent requests, and roughly 25% faster single requests. This isn't just about Macs running AI faster; it's that AI can now write sufficiently good low-level performance code.
What this is
Splash is the key software that lets large models run locally on Mac (similar to a "driver"), and is one of the faster engines in the local inference (running AI models on your own computer without relying on cloud servers) ecosystem. The developer forked Splash (meaning "copy the original code and modify it yourself"), and did four things targeting M5 Max's 40-core CPU:
First, used Splash's built-in tuner (automatic parameter-tuning tool) to select the optimal compute path for M5 Max, boosting single-request speed by about 20%. Second, wrote new verify kernels (low-level programs that verify computation correctness) for M5's tensor unit (matrix-computation accelerator, the core hardware for AI inference), gaining 10–19% in concurrent scenarios. Third, tuned for the 35B-parameter class Qwen3.6 model, with concurrent speedups of 18–22%. Fourth, borrowed the TensorFold approach, letting the model directly copy original segments when rewriting known text, achieving up to 42% faster full-file edits.
In actual testing, generation speed went from 45–51 tokens per second (the smallest unit AI processes text, roughly a "character") to 56–64, with quality unchanged and bit-identical results.
Industry view
The supportive view: the ceiling for local inference has been pushed up again. Mid-sized companies, research institutions, and data-sensitive finance/healthcare/legal teams used to either pay high prices for cloud-vendor compute, or tolerate slow local performance — now there's an additional option of "Mac workstation + optimized engine."
There are plenty of voices urging caution:
First, M5 Max has not been officially released as of publication; the data comes from a developer's self-testing on a test machine, and cannot simply be equated with real consumer experience. Second, engines like Splash depend heavily on community maintenance; whether developers are willing to keep pace with the M5 series long-term is a question — a single person's fork has no SLA. Third, the speed you can get from running 4-bit quantization (compressing the model to one-quarter size at the cost of reduced precision) is questionable for enterprise production environments; many scenarios require both throughput and precision.
What's more worth discussing is this developer's working method: performance tuning, writing code, and running benchmarks in one shot, delivered in days. AI's ability to write performance-sensitive code has passed the "it works" stage and entered the "it can optimize" stage.
Impact on regular people
- For enterprise IT: hardware costs for private deployment of large models may continue to drop, but compatibility and long-term maintenance risks of Mac clusters need to be assessed.
- For working professionals: running local AI assistants on Mac for tasks like copywriting, coding, and reporting will get closer to the cloud experience, without needing to go online every time.
- For the consumer market: consumer PCs with NPUs (Neural Processing Units, dedicated AI accelerator chips) will become the new standard, but short-term impact on most people is limited — it's mainly prepared for "people who frequently run models."