local inference
13 articles tagged with this topic
llama.cpp Has 50 PRs Pending — Local AI No Longer Needs a High-End GPU
Open-source llama.cpp has 50+ performance PRs pending merge, some claiming 3x CPU inference speedup. Local LLM deployment is shedding its dependence o
Zhipu GLM Runs Locally on Mac — And This Matters More Than It Looks
ds4 (co-maintained by Redis creator antirez) added Zhipu GLM Flash support this week, running on 128GB M4 Max. A concrete step for local Chinese LLMs.
Qwen3 Hits 220 tokens/sec on 5090 — Local AI Inflection Point Nears
New inference engine Ninfer pushed Qwen3 to 170 avg / 220 peak tokens/sec on RTX 5090 — over 2x faster than llama.cpp. Local AI is closing in on cloud
Local Hack Triples 256k Inference Speed, but Commercial Hurdles Remain
Reddit user validates chunked KV cache on a small open-source model: 256k prefill runs 3x faster, needle-in-haystack accuracy holds. Real value: bottl
33B Audio-Video Model on Apple Silicon: Open Source Tallies Acceleration's True Cost
This week, h3.c ported a 33B audio-video model to Apple Silicon, exposing six distinct optimization knobs instead of a single "fast=true" switch.
Laptop + eGPU box delivers 40GB VRAM, local Qwen3 27B runs 70% faster
Reddit user pairs laptop + eGPU box via Thunderbolt 4 to hit 40GB VRAM for local Qwen3 27B, boosting speed from 16 to 27 tokens/s (≈70%).
Qwen 27B Hits 138 Tokens/Sec on a Single RTX 3090 — Local AI Costs Crater
Qwen 27B hits 138 tokens/sec on one RTX 3090; 2nd-turn latency from 23s to 1s. Not a model breakthrough — open-source is flattening local LLM costs.
Liquid AI Ships 2.6B Local Model — The Small-Model Path Is Getting Real
Liquid AI releases 2.6B-parameter LFM 2.5 in GGUF format, runnable on laptops. Paired with Phi, Llama, and Qwen, "local AI" is quietly becoming an ent
Ling-3.0 Lands Official llama.cpp Support — 128K Context on 12GB GPUs
China's open-source Ling-3.0 joins llama.cpp main branch — full 128K context on consumer 12GB GPUs at ~110 tokens/sec. Private deployment just got che
4 Consumer GPUs Run a 144GB LLM — Local AI's Cost Inflection Point Is Here
A developer ran a 144GB quantized DeepSeek-V4-Flash on 4 RTX 3060s + a standard workstation, hitting ~100 token/s. Local LLMs may no longer need A100/
124B model ran stably 7 min on desktop — local AI crosses usability threshold
Reddit test: 124B Ling model held 35.7 tok/s for 15K+ tokens on a single NVIDIA DGX Spark — local LLMs shifting from geek toy to enterprise option.
Meta Crams AI Models Into Your Phone — But Is On-Device Really Worth It?
Meta open-sourced two phone-runnable models, Muse Glimmer and Muse Spark 1.2. We ask: is local inference actually cheaper than cloud APIs?
AMD In-House AI Mini PC in June: Chipmaker Building Systems is a Major Signal
AMD's in-house Ryzen AI 395 mini PC (June, Lenovo OEM) shows local AI inference moving from concept to product as chipmakers pivot from parts to syste