One number: 50 tokens/second—this week the open-source Splash 1.1.0 delivered smooth local inference of a 27-billion-parameter large model on Apple Silicon. The implication: enterprises no longer have to pay cloud fees for every AI call.

What this is

Splash is a local AI inference engine for Apple Silicon (Apple's in-house M-series chips) developed by the incoai team. This week's 1.1.0 release adds support for GGUF (an open-source model compression format) and MLX (Apple's official machine learning framework).

Reddit user benchmarks show: running a Qwen3-series 27-billion-parameter model (parameters roughly map to a model's "brain capacity"; 27B sits in the upper-mid range) on a 64GB M5 Pro hits roughly 50 tokens/second—about the pace of fluent human reading.

Industry view

Supporters call this a "breakthrough" for Apple Silicon local inference. One user wrote it's the first time an agentic setup (AI autonomously completing multi-step tasks) feels genuinely usable on a Mac.

But we note several reservations:

  • The benchmarked machine is a top-spec 64GB M5 Pro that costs 20,000+ RMB—out of reach for most users
  • A 27B model still answers weaker than cloud-side 100B+ flagship models
  • The project is early-stage with a small community; long-term maintenance is uncertain
  • Local inference remains the domain of enthusiast tinkerers; mainstream enterprise IT hasn't really engaged

Impact on regular people

For enterprise IT: sensitive data (contracts, customer records, internal documents) now has an AI processing path that never leaves the company. Finance, healthcare, and legal can reassess whether every call must go to the cloud.

For individual professionals: knowledge workers may soon run personal AI assistants directly on company-issued Macs, with significantly better availability for travel and offline scenarios.

For the consumer market: once local AI goes mainstream, a 200 RMB/month ChatGPT subscription may no longer be mandatory—and hardware vendors like Apple could emerge as new winners.