Back to home

MLX

7 articles tagged with this topic

AppleMac

Mac Local LLM Inference: A Complete Mess, a Full Generation Behind NVIDIA

A Reddit developer tested all major Mac LLM frameworks for two weeks. Verdict: Apple's AI software is fragmented, a generation behind NVIDIA.

Aug 162 min read
MLXLocalLLaMA

MLX 4bit Quantization Showdown: Which Compression Format Actually Wins on Apple Silicon?

A Reddit thread compares four 4bit quantization schemes for running Qwen3.6 on Apple silicon. We break down what each format trades off — and why it m

Aug 82 min read
OllamaQwen

Ollama Runs Local LLMs on Mac with One Command — PCs Are the New AI Gateway

Ollama runs Qwen & DeepSeek locally on Mac via one command. MLX integration doubles inference speed. When deployment = app install, cloud-free AI may

May 22 min read
MLXQwen3.5

DFlash speculative decoding on Apple Silicon: 4.1x on Qwen3.5-9B, now open source (MLX, M5 Max)

Open-source DFlash achiev es 4.13x speedup on Qwen3.5-9B using MLX on M5 Max with 89.4% token acceptance rate.

Apr 132 min read
CoreMLApple-Intelligence

Apple's On-Device AI Moat: What It Means for Edge Builders

Apple's privacy-first, on-device AI stack may become the default for builders who need inference without cloud costs .

Apr 132 min read
Gemma 4MLX

Gemma 4 audio with MLX

Google's Gemma 4 E2B model can transcribe audio locally on macOS using MLX and a single uv run command.

Apr 132 min read
HitokuGemma-4

Hitoku, open-source local macOS context aware assistant with Qwen3.5/Gemma4

Open-source macOS assistant runs Gemma 4 and Qwen 3.5 fully on-device with screen and document context .

Apr 132 min read