Apple Silicon
8 articles tagged with this topic
Qwen 3.8 Tested: Deep Thinking Burns 5.5x Tokens—Local Deployment Math Changes
Reddit user tested Qwen3.8-27B on M5 Max: deep thinking uses 5.5x tokens, 6x time; disabling tanks quality. The "thinking" cost gap is exposed.
Mac mini's M6 upgrade drops the bar for running AI models locally
Apple's Mac mini refresh with next-gen Apple Silicon signals local AI is moving from geek toy toward semi-mainstream — but production readiness still
33B Audio-Video Model on Apple Silicon: Open Source Tallies Acceleration's True Cost
This week, h3.c ported a 33B audio-video model to Apple Silicon, exposing six distinct optimization knobs instead of a single "fast=true" switch.
Qwen Runs 45 Tokens/Second Locally — Apple Silicon's Silent Win for Open LLMs
Reddit user benchmarked Qwen 3.8B at 45+ tok/s on Apple silicon — a record on consumer hardware. Local LLM inference shifts from geek toy to viable op
Apple Silicon Runs Qwen 27B 3x Faster — Local LLMs Enter Usable Territory
mlx-dspark ports DeepSeek's speculative decoding to Apple Silicon, giving Qwen 27B a 3x speedup on M-series Macs with no quality loss. Local LLMs near
Meta's Muse Glimmer 30B Hits 3.3x Speedup on Mac as Open Source Closes the Gap
Developer A-Rahim uses speculative decoding to push Meta's Muse Glimmer 30B up to 3.3x faster on M4 Pro, with output exactly matching the original.
Local AI Gets Serious: Anubis-OSS Leaderboard Tracks 218 Models, 10 Apple Chips
Anubis-OSS leaderboard updates: 371 submissions, 218 models, 10 Apple chips. This data proves local open-source model deployment is no longer a geek t
37 LLMs Benchmarked on MacBook Air M5 32GB: Full Speed Results
Community benchmark of 37 local LLMs on M5 Air 32GB using llama-bench reveals MoE models as clear winners for speed-to-quality ratio.