Back to home

local inference

13 articles tagged with this topic

llama.cppMoE

llama.cpp Has 50 PRs Pending — Local AI No Longer Needs a High-End GPU

Open-source llama.cpp has 50+ performance PRs pending merge, some claiming 3x CPU inference speedup. Local LLM deployment is shedding its dependence o

4h ago2 min read
ZhipuGLM

Zhipu GLM Runs Locally on Mac — And This Matters More Than It Looks

ds4 (co-maintained by Redis creator antirez) added Zhipu GLM Flash support this week, running on 128GB M4 Max. A concrete step for local Chinese LLMs.

1d ago2 min read
NinferQwen3

Qwen3 Hits 220 tokens/sec on 5090 — Local AI Inflection Point Nears

New inference engine Ninfer pushed Qwen3 to 170 avg / 220 peak tokens/sec on RTX 5090 — over 2x faster than llama.cpp. Local AI is closing in on cloud

1d ago2 min read
KV cacheLocalLLaMA

Local Hack Triples 256k Inference Speed, but Commercial Hurdles Remain

Reddit user validates chunked KV cache on a small open-source model: 256k prefill runs 3x faster, needle-in-haystack accuracy holds. Real value: bottl

Aug 222 min read
h3.cApple Silicon

33B Audio-Video Model on Apple Silicon: Open Source Tallies Acceleration's True Cost

This week, h3.c ported a 33B audio-video model to Apple Silicon, exposing six distinct optimization knobs instead of a single "fast=true" switch.

Aug 222 min read
Qwen3llama.cpp

Laptop + eGPU box delivers 40GB VRAM, local Qwen3 27B runs 70% faster

Reddit user pairs laptop + eGPU box via Thunderbolt 4 to hit 40GB VRAM for local Qwen3 27B, boosting speed from 16 to 27 tokens/s (≈70%).

Aug 202 min read
QwenvLLM

Qwen 27B Hits 138 Tokens/Sec on a Single RTX 3090 — Local AI Costs Crater

Qwen 27B hits 138 tokens/sec on one RTX 3090; 2nd-turn latency from 23s to 1s. Not a model breakthrough — open-source is flattening local LLM costs.

Aug 192 min read
Liquid AILFM 2.5

Liquid AI Ships 2.6B Local Model — The Small-Model Path Is Getting Real

Liquid AI releases 2.6B-parameter LFM 2.5 in GGUF format, runnable on laptops. Paired with Phi, Llama, and Qwen, "local AI" is quietly becoming an ent

Aug 192 min read
Ling-3.0llama.cpp

Ling-3.0 Lands Official llama.cpp Support — 128K Context on 12GB GPUs

China's open-source Ling-3.0 joins llama.cpp main branch — full 128K context on consumer 12GB GPUs at ~110 tokens/sec. Private deployment just got che

Aug 182 min read
DeepSeekRTX 3060

4 Consumer GPUs Run a 144GB LLM — Local AI's Cost Inflection Point Is Here

A developer ran a 144GB quantized DeepSeek-V4-Flash on 4 RTX 3060s + a standard workstation, hitting ~100 token/s. Local LLMs may no longer need A100/

Aug 182 min read
inclusionAILing

124B model ran stably 7 min on desktop — local AI crosses usability threshold

Reddit test: 124B Ling model held 35.7 tok/s for 15K+ tokens on a single NVIDIA DGX Spark — local LLMs shifting from geek toy to enterprise option.

Aug 132 min read
MetaMuse Glimmer

Meta Crams AI Models Into Your Phone — But Is On-Device Really Worth It?

Meta open-sourced two phone-runnable models, Muse Glimmer and Muse Spark 1.2. We ask: is local inference actually cheaper than cloud APIs?

Aug 102 min read
AMDRyzen AI 395

AMD In-House AI Mini PC in June: Chipmaker Building Systems is a Major Signal

AMD's in-house Ryzen AI 395 mini PC (June, Lenovo OEM) shows local AI inference moving from concept to product as chipmakers pivot from parts to syste

Apr 302 min read