Back to home

local deployment

30 articles tagged with this topic

DeepSeekGLM

GLM Beats DeepSeek on Two GPUs — Chinese Open-Source Stops Compromising

On two NVIDIA DGX Sparks, GLM-5.3 Flash beat DeepSeek V4 Flash on HumanEval (97% vs 94.5%). GLM ran 30% slower with one-quarter the context.

Just now2 min read
QwenAlibaba

Engineer Pushes Qwen to the Limit: 260K Tokens Is Local AI's Hard Ceiling

Engineer pushed Qwen to extremes: context over 100K tokens drops generation 75%. Long-context remains local AI's hard ceiling—proof enterprises can't

2h ago2 min read
DeepSeekNVIDIA

DeepSeek Hits 67 token/s on Two $9K Mini Boxes — Local LLM Floor Is Caving In

Reddit user hit 67-84 token/s on DeepSeek V4 Flash with a 1M-token context window on two ~$9K NVIDIA DGX Sparks. The local-LLM cost barrier is collaps

4h ago2 min read
QwenAlibaba

Qwen 27B Runs on Just 18GB VRAM — Local LLM Bar Drops Again

Qwen's 27B multimodal needs only 18GB VRAM (Q4), runnable on consumer GPUs. Local LLM bar drops again, but MoE version still demands clusters.

1d ago2 min read
ZhipuTensorSharp

TensorSharp Doubles Llama.cpp Decoding Speed, Could Halve LLM Deployment Costs

GLM-5.3-Flash hits 2x llama.cpp decoding on new TensorSharp framework; local deployment costs may drop sharply, but ecosystem maturity is unproven.

1d ago2 min read
QwenAlibaba

Qwen Runs Locally — The "AI-Must-Be-Cloud" Assumption Cracks

Alibaba's Qwen3.8-Flash-Next lands in llama.cpp, letting regular PCs run AI locally without cloud APIs—open-source's quiet win over closed vendors.

2d ago2 min read
QwenNVIDIA

Lunchbox Rig Runs 27B AI Model — Local Players Catch Up to Paid Services

A Reddit rig pairs a Toughbook with an RTX Pro 6000 to run a 27B model at 262K context — and beats Gemini Pro and ChatGPT on legal OCR.

4d ago2 min read
Moonshot AIDeepSeek

2.8T-Parameter Model Crammed Into Gaming PC — Runs, But Too Slow to Use

Engineer runs 711GB, 2.8T-param Kimi K3 on RTX 4070 Ti + 32GB RAM — normal speed 1-2 tok/s. Dual SSD parallelism pushes DeepSeek from 1.14 to 8.08 tok

4d ago2 min read
QwenNVIDIA

Local AI Is Finally Competitive — What 30B Open-Source Benchmarks Reveal

LocalLLaMA benchmark shows 30B open-source models (Ornith, TielCoder, Qwen, Nemotron) now rival closed APIs on coding—local AI deployment becomes viab

5d ago2 min read
Qwenopen-source model

16GB VRAM Hits Flagship Performance — Closed-Source Giants Go Quiet

Qwen's new 27B open-source model runs locally on a single 16GB gaming GPU at near-flagship quality. Notably, OpenAI and Anthropic stayed silent this t

Aug 232 min read
QwenAlibaba

Qwen 3.8-27B One Week In: Top Marks for Doing, Memory Slips

Qwen 3.8-27B earned 'local best' on agent tasks in 2,000 Reddit tests but regressed on knowledge memory—first open-source model approaching GPT on exe

Aug 232 min read
AlibabaQwen3.8

Alibaba packs Opus-class AI into 24GB GPUs; open-source local LLMs truly work

Qwen3.8-27B hit Hugging Face trending #1 in 48 hours; Cline made it default in 4 days. Open-source LLMs hit a usable threshold for the first time.

Aug 222 min read
Llama.cppMeta

Llama.cpp Hits 0.2 — The Bar for Running LLMs on Home PCs Drops Again

Open-source inference engine Llama.cpp releases 0.2.0, its first 0.2-series version, signaling a systemic architecture and performance overhaul worth

Aug 222 min read
AlibabaQwen

Alibaba Qwen Ships Twice in 6 Months — Local-Usable Builds Always Lag 2 Months

Alibaba's Qwen shipped two versions in six months, but the locally deployable versions always trail the benchmark leader by two months.

Aug 222 min read
llama.cppQwen3

Reddit 用户让本地 AI 提速 65%,但官方还没接盘

llama.cpp fork with DSpark PC Tree speculative decoding pushes Qwen3 ~65% faster on RTX 5090. Not merged, but local AI is getting cheaper and faster.

Aug 212 min read
OrnithNvidia 5090

Ornith-1.5 hits 250 token/s on Nvidia 5090 — local AI nears cloud parity

Ornith-1.5-35B-A3B open-source model hits 250 token/s on Nvidia 5090 — first realistic local option for cloud-grade Agent workloads.

Aug 212 min read
QwenAMD

Consumer GPU Hits 153 tok/s on a 27B Model — Local AI Cost Inflection Arrives

A Reddit developer ran 159 experiments: Qwen3.8-27B on consumer hardware beats dual-3090 cloud servers on HumanEval. Just swapping the chat template s

Aug 212 min read
QwenTongyi Qwen

Qwen 27B on RTX 3090s draws nested SVG; mid-size open-source models underrated

Qwen 27B drew nested SVG on two RTX 3090s at 46k tokens. Mid-size open-source creativity on consumer hardware looks underrated; ~$10K GPUs stay a barr

Aug 202 min read
UnslothQwen

Unsloth squeezes Qwen onto 8GB laptops — local LLMs now run on almost anything

Unsloth's Dynamic v3.0 lets Qwen models run on 8GB laptops, with 1-bit versions keeping 77% accuracy — easing local enterprise LLM deployment.

Aug 192 min read
AlibabaQwen

Qwen 3.8 27B Halves Agent Coding Requests, Catches DeepSeek in Open-Source Race

Reddit r/LocalLLaMA benchmark: Qwen 3.8 27B matches DeepSeek Flash in agent coding, halves requests vs 3.6, runs on one GPU.

Aug 182 min read
Qwenopen-source LLM

Two Prompts, One Game — But Local Qwen's Barrier Isn't as Low as You Think

Reddit user built a web game with two prompts on local quantized Qwen. Open-source LLMs now handle practical tasks, but hardware barriers remain.

Aug 182 min read
AlibabaQwen3

Alibaba's Qwen3 27B Compressed to 18GB, Runs on Single GPU — Local LLM Bar Drops Again

Alibaba's Qwen3 27B compressed to 18GB via int4 quantization with MTP acceleration, runs on a single consumer GPU. Local LLM hardware costs keep falli

Aug 162 min read
Qwenopen-source models

Local Open-Source LLMs Booming for Coding — The Real Battle Is Now the Harness

Reddit poll this week: developers now pick the harness that wires AI into their code workflow. The fight is moving from the model to the tool layer.

Aug 162 min read
QwenAlibaba

Qwen 27B Hits 3.8 — China Open-Source LLMs Keep the Local-Deployment Race Hot

Alibaba's Qwen shipped 27B v3.8. The real story isn't the version bump but changed sampling parameters—developers must retune or see output drift.

Aug 152 min read
QwenAlibaba

Qwen3.8 Local Test: One Reasoning Dial, 20x Token Cost — The Hidden Bill

Qwen3.8-27B local test: reasoning_effort swings tokens 20x (2K–40K). "Thinking cost" is the underestimated deployment variable in the reasoning era.

Aug 142 min read
QwenAlibaba

Qwen 27B Sparks Overseas Buzz — China's Open-Source LLMs Step Into the Ring

Alibaba ships Qwen 27B to overseas devs for VRAM and speed tests — a concrete signal China's open-source LLMs now compete head-on with Llama and DeepS

Aug 142 min read
QwenTongyi Qianwen

Qwen3.8-27B Drops Early — 27B Is Local AI's Sweet Spot, Benchmarks Pending

Alibaba's Qwen team posted the Qwen3.8-27B model card on Hugging Face ahead of benchmarks. 27B is open-source's local AI sweet spot.

Aug 142 min read
UnslothMuse-Glimmer

Unsloth Releases 30B Open-Source Model — Local LLMs Step Out of the Geek Bubble

Unsloth ships a GGUF-quantized Muse-Glimmer-30B, letting consumer laptops run a 30-billion-parameter LLM locally. On-prem AI is shifting from hobbyist

Aug 102 min read
LocalLLaMAopen source models

A Developer Wants a Local Model to Turn Text Into Regex — And That Tiny Request Hides a Real

A Reddit post asking for local small models that convert natural language to regex reveals a real enterprise pain point: extracting structured info fr

Aug 92 min read
DeepSeekV4-Flash

DeepSeek V4-Flash Falls Asleep Mid-Task: The Hidden Cost of Local AI Deployment

DeepSeek's V4-Flash open-source model silently halts generation past 100k tokens. Concrete proof that "downloadable" and "reliably usable" are still f

Aug 92 min read