Back to home
Splash
3 articles tagged with this topic
AppleM5 Max
Mac Local LLMs Hit 50% Faster — AI Wrote the Performance Code
Developer used Claude to rewrite Splash inference engine for M5 Max: 50% faster concurrent, ~25% single request. AI now writes performance code.
Sep 272 min read
SplashApple Silicon
Splash 1.1: Apple Silicon Now Runs 27B AI Models Smoothly
Splash 1.1.0 runs 27B models locally on top-spec Macs at ~50 tokens/sec. Local AI inference becomes practical—enterprises can skip cloud fees.
Sep 262 min read
AppleM5 Max
M5 Max Mac Local AI 20% Faster—Hidden Cost: Per-Gen Kernel Retuning Required
Developer retunes inference kernels for 40-core M5 Max, boosting local LLM generation 20%. Each Apple Silicon generation needs manual AI re-optimizati
Sep 252 min read