MLX
7 articles tagged with this topic
Mac Local LLM Inference: A Complete Mess, a Full Generation Behind NVIDIA
A Reddit developer tested all major Mac LLM frameworks for two weeks. Verdict: Apple's AI software is fragmented, a generation behind NVIDIA.
MLX 4bit Quantization Showdown: Which Compression Format Actually Wins on Apple Silicon?
A Reddit thread compares four 4bit quantization schemes for running Qwen3.6 on Apple silicon. We break down what each format trades off — and why it m
Ollama Runs Local LLMs on Mac with One Command — PCs Are the New AI Gateway
Ollama runs Qwen & DeepSeek locally on Mac via one command. MLX integration doubles inference speed. When deployment = app install, cloud-free AI may
DFlash speculative decoding on Apple Silicon: 4.1x on Qwen3.5-9B, now open source (MLX, M5 Max)
Open-source DFlash achiev es 4.13x speedup on Qwen3.5-9B using MLX on M5 Max with 89.4% token acceptance rate.
Apple's On-Device AI Moat: What It Means for Edge Builders
Apple's privacy-first, on-device AI stack may become the default for builders who need inference without cloud costs .
Gemma 4 audio with MLX
Google's Gemma 4 E2B model can transcribe audio locally on macOS using MLX and a single uv run command.
Hitoku, open-source local macOS context aware assistant with Qwen3.5/Gemma4
Open-source macOS assistant runs Gemma 4 and Qwen 3.5 fully on-device with screen and document context .