Back to home

local LLM

17 articles tagged with this topic

AMDllama.cpp

AMD GPUs Now 30% Faster on Local LLMs — Another Reason to Escape NVIDIA

llama.cpp's new Vulkan backend for AMD RDNA3/RDNA4 GPUs delivers ~27% throughput boost on 26B models. A signal for SMEs eyeing local AI without NVIDIA

Sep 242 min read
QwenRTX 4090

RTX 4090 Runs Qwen Locally — But Local AI Isn't Ready for White-Collar Work

A Reddit user ran Qwen 27B on an RTX 4090 for coding. Chinese open-source gains traction with developers; local LLMs remain a hobbyist toy.

Sep 232 min read
QwenAlibaba

Qwen 1-bit is still 6x slower — companies eyeing local LLMs should wait

Qwen's 1-bit quantization runs 6x slower than the 4-bit version with ~70% accuracy. Local LLM deployment isn't ready to replace cloud APIs.

Aug 302 min read
Qwenlocal LLM

Local AI Coding Is Trending — But Most Companies' GPUs Can't Run It

A Reddit post about running Qwen 3 27B locally on an RTX A4500 for AI coding sparked debate. Local model coding is shifting from hobbyist toy to real

Aug 302 min read
QwenAlibaba Tongyi

Community Qwen3.8 Quantization Saves 30GB — Local LLM Bar Drops Again

Community dev agentionai ships a custom quantized Qwen3.8-Flash-Next, 20–30GB smaller than mainstream versions at comparable quality. Local LLM bar dr

Aug 292 min read
llama.cpplocal LLM

llama.cpp Adds 'Lazy Loading': 100B Models Run Locally, Hardware Bar Drops

llama.cpp adds TENSOR_READ_LAZY: weights load on-demand instead of fully residing in memory. Enables larger models on consumer hardware—a real shift f

Aug 272 min read
DFlash2llama.cpp

DFlash2 Makes Qwen 27B Inference 3x Faster — Local Mid-Sized LLMs Now Workable

DFlash2, integrated into llama.cpp, makes Qwen3 27B generation 3x faster. Local mid-sized LLMs shift from hobbyist tinkering to viable tooling.

Aug 192 min read
AlibabaQwen

Alibaba's Qwen Local Updates Again — Another Option for Running LLMs on Your PC

Alibaba's Qwen local GGUF build updates with lower VRAM needs. The bar for self-hosted LLMs drops a notch, but this is routine open-source maintenance

Aug 192 min read
Redstartllama.cpp

One dev crammed enterprise AI stack into a PC — local LLMs just got practical

Indie dev Redstart packs MCP, tool permissions, and model hosting into one PC — local LLMs cross from hobbyist to usable.

Aug 162 min read
DeepSeeklocal LLM

DeepSeek Flash 0731 Hits Flagship Scores on a $2K Laptop — Cloud's Edge Erodes

DeepSeek's Flash 0731 scores flagship on Artificial Analysis. A Reddit user reports it runs on a 2021 sub-$2000 consumer laptop — local AI's tipping p

Aug 142 min read
QwenAlibaba

1500 Yuan GPU Runs Qwen 30B — Local LLM Hardware Bar Drops Again

Reddit user runs Qwen 30B MoE on RTX 3050 6GB, hitting 30 tokens/sec at 90k context. Enterprise IT should reassess local LLM costs.

Aug 142 min read
MiniMaxMusic3

MiniMax-Music3 Launches to Zero Buzz — Music AI Race Is Too Crowded

MiniMax-Music3 hit r/LocalLLaMA with just a title—no specs, no demo, no comparisons, zero comments. We can't yet tell real iteration from hype.

Aug 132 min read
QwenTongyi Qianwen

Qwen 3.5 Hits 18 token/s on a $280 Radeon 7600 — Local LLMs Are Finally "Good Enough"

Alibaba's Qwen 3.5 35B MoE model runs at 18 token/s on a ~$280 Radeon 7600 via llama.cpp, pushing local LLMs into genuinely usable territory.

Aug 102 min read
RTX 5090AMD 9700

Three RTX 5090s Aren't Enough: How Long Can the Local LLM Hardware Arms Race Burn?

A user with three RTX 5090s debates adding an AMD AI Pro to hit 128GB VRAM for DeepSeek. Local LLM deployment is shifting from "can it run" to "how bi

Aug 92 min read
QwenAMD

Qwen3.5 Runs Fast on AMD GPUs, Slow on NVIDIA — A Bug Nobody Understands Yet

The same Qwen3.5-9B model runs 32% slower than baseline on NVIDIA via CUDA, but 77%-128% faster on AMD via Vulkan. One unverified report, but local-LL

Aug 82 min read
Qwen3local LLM

本地运行 AI 编程时, 要不要关掉「思考模式」?一个值得厘 清的实用问题

Should you disable thinking mode when running Qwen3 locally for coding? A real debate with structural implications for AI dev toolch ains.

Apr 182 min read
llama.cppGemma 4

Gemma 4 llama.cpp Issues Resolved With Recent Fixes

Google Gemma 4 models now run correctly in llama.cpp after critical fixes for output quality and crashes

Apr 42 min read