Back to home

Qwen

30 articles tagged with this topic

QwenAlibaba

Qwen 27B on a Pi Dumps Cloud APIs — Open-Source AI Is Finally 'Good Enough'

Dev runs Qwen 27B on Pi, abandons OpenAI and Anthropic APIs. Local cost matches cheapest cloud—a reset signal for open-source AI's commercial value.

Sep 242 min read
DeepSeekQwen

Deploying LLMs Means Tuning 40+ Parameters—It's Becoming an Independent Craft

40+ parameters to tune per LLM deployment: a new guide reveals AI production has hit engineering hard reality—tuning is now a standalone craft.

Sep 242 min read
AliyunCloud Geography

Aliyun Expands to Turkey, Finland, Netherlands — Export Edge Isn't the Model

Aliyun adds data centers in Turkey, Finland, Netherlands. The real AI-export bottleneck isn't models—it's cloud geography, the compliance and cost var

Sep 242 min read
DewuQwen

Dewu Rebuilds Search With LLMs — But Generative Recall's Cost Is Unclear

Dewu's 'generative recall' rebuilds search with LLM-generated tags—the first systematic engineering writeup from a major Chinese e-commerce team.

Sep 242 min read
AppleMac Studio

$7K Mac Studio Hits 1,700 Agent Schedules — Local LLMs Finally Do Real Work

Reddit dev runs Qwen 3 on M5 Ultra Mac Studio (96GB), hits 1,700 sub-agent schedules and 89% cache hit on 112M-token dev task — first engineering-grad

Sep 232 min read
LocalLLaMALlama.cpp

40MB Single File Packs LLM + Voice — Local-First AI Goes Live

Indie dev packed LLM, voice, code execution into a 40MB local exe; AI generates every UI pixel live. Rare working Local-First demo challenging cloud v

Sep 232 min read
QwenTongyi Qianwen

Qwen 27B Now Runs on One AMD GPU — Local LLMs Start Feeling Installable

Reddit dev ships Qwen 27B GGUF quantization on a single AMD Strix Halo, beating ISTA and Unsloth on three benchmarks. One person, one GPU, one week.

Sep 232 min read
QwenRTX 4090

RTX 4090 Runs Qwen Locally — But Local AI Isn't Ready for White-Collar Work

A Reddit user ran Qwen 27B on an RTX 4090 for coding. Chinese open-source gains traction with developers; local LLMs remain a hobbyist toy.

Sep 232 min read
QwenAlibaba

Overseas Devs Build Qwen Coding Tools — Alibaba's Open-Source Bet Pays Off

Reddit's r/LocalLLaMA debated which 'harness' fits Qwen 3.8 best. Surface-level tool fight; underneath, Qwen is now a Western open-source default.

Sep 232 min read
Redditopen-webui

Reddit User Deploys AI Agent at Home — Early Local LLM Experiments

A popular Reddit post shows a home AI Agent setup: open-webui + isolated VM + Qwen LLM. We examine what it reveals about local LLMs leaving data cente

Sep 232 min read
QwenTongyi Qianwen

Apple M5 Max hits 350K-token context — hardware bar cleared, quality still lags

128GB MacBook Pro holds Qwen3.8 at 350K tokens, but past 100K the model 'confuses roles.' Hardware bar falls; quality still lags.

Aug 302 min read
QwenAlibaba

Qwen 1-bit is still 6x slower — companies eyeing local LLMs should wait

Qwen's 1-bit quantization runs 6x slower than the 4-bit version with ~70% accuracy. Local LLM deployment isn't ready to replace cloud APIs.

Aug 302 min read
QwenAlibaba

Engineer Pushes Qwen to the Limit: 260K Tokens Is Local AI's Hard Ceiling

Engineer pushed Qwen to extremes: context over 100K tokens drops generation 75%. Long-context remains local AI's hard ceiling—proof enterprises can't

Aug 302 min read
Qwenlocal LLM

Local AI Coding Is Trending — But Most Companies' GPUs Can't Run It

A Reddit post about running Qwen 3 27B locally on an RTX A4500 for AI coding sparked debate. Local model coding is shifting from hobbyist toy to real

Aug 302 min read
NInfervLLM

Million-Token Context on Two 5090s — Amateur Dev Shatters Enterprise AI Myth

Reddit developer NInfer hits 1.04M token context on consumer RTX 5090s at 119 tok/s — 2.8x faster than vLLM with a 27B Qwen model.

Aug 302 min read
QwenApple Silicon

Qwen 3.8 Tested: Deep Thinking Burns 5.5x Tokens—Local Deployment Math Changes

Reddit user tested Qwen3.8-27B on M5 Max: deep thinking uses 5.5x tokens, 6x time; disabling tanks quality. The "thinking" cost gap is exposed.

Aug 292 min read
QwenRTX 5080

RTX 5080 Hits Just 6 Characters/Second on Qwen — How High Is Local AI's Home Threshold?

Reddit user hit ~6 chars/sec running quantized Qwen on RTX 5080 + 64GB RAM. Local AI still demands far more hardware and tuning than ordinary users ca

Aug 292 min read
QwenAlibaba Tongyi

Community Qwen3.8 Quantization Saves 30GB — Local LLM Bar Drops Again

Community dev agentionai ships a custom quantized Qwen3.8-Flash-Next, 20–30GB smaller than mainstream versions at comparable quality. Local LLM bar dr

Aug 292 min read
QwenTongyi Qianwen

Qwen 27B hits 50 tok/s: 16GB consumer GPUs can now run large models locally

A Reddit user ran Alibaba's Qwen 27B on a 16GB consumer GPU, hitting 50 tok/s generation + 100k context. Local LLMs are leaving the geek circle behind

Aug 292 min read
NVIDIADGX Spark

Maxing AI 'Thinking Depth' Hurts Results — DGX Spark Local Test Warns Enterprises

A Reddit developer tested DeepSeek/Qwen on four DGX Sparks: 'deep thinking' mode lowers scores and doubles runtime — a direct cost warning for AI infe

Aug 292 min read
QwenAlibaba

Qwen Makes Thinking Depth Adjustable — Alibaba Lets LLMs Allocate Compute On Demand

Alibaba's Qwen now lets users adjust 'thinking depth'—quick answers for easy questions, more reasoning for hard ones. LLMs shift from on/off switch to

Aug 292 min read
QwenGLM

Alibaba and Zhipu Bet on Small Models — Local AI Faces Choice Overload

Qwen Flash and GLM Flash launched together, leaving local users with choice overload. China's open-source LLMs shift from parameter wars to same-tier

Aug 292 min read
QwenAlibaba

Alibaba Open-Source Model Revives a Bricked Foldable — Local AI Runs Solo

Reddit user revived a bricked foldable with Qwen 3.8 (27B) on a Raspberry Pi, saving ~$600. Real signal: local AI trusted solo on critical tasks.

Aug 292 min read
QwenAlibaba

Qwen 27B Runs on Just 18GB VRAM — Local LLM Bar Drops Again

Qwen's 27B multimodal needs only 18GB VRAM (Q4), runnable on consumer GPUs. Local LLM bar drops again, but MoE version still demands clusters.

Aug 292 min read
QwenQuantization

Alibaba Qwen 27B Squeezed to 10GB, Matches Original Quality — Local AI Advances

Austria's ISTA-DASLab squeezed a Qwen 27B to 10GB, matching original quality. Local AI advances — capable, private deployment without cloud uploads.

Aug 292 min read
QwenNVIDIA

Two DGX Sparks Hit 181 tok/s Concurrent: Local Multi-Agent Is Now Viable

Two NVIDIA DGX Sparks + Qwen models hit 181 tok/s concurrent across 9 parallel AI Agents. Local multi-Agent moves from demo to real work.

Aug 292 min read
ZhipuGLM-5.3

Zhipu GLM-5.3 Scores Higher Without New Architecture — Chinese LLMs Turn Inward

Zhipu's GLM-5.3 hits stronger benchmarks with architecture identical to 5.2 — a training-only upgrade worth watching while peers chase new designs.

Aug 282 min read
QwenRTX 3090

Two 3090s Run 27B Model at 165 tok/s — Local AI Is Finally 'Good Enough'

We noted a Reddit user hit 165 tok/s on a 27B Qwen model using two RTX 3090s (used rig under ¥20K) — Agent-ready. Local LLMs just crossed from 'toy' t

Aug 282 min read
QwenLocalLLaMA

Local 8B Model Ties 27B on Agent Coding — The 'Good Enough' Moment Arrives

Reddit dev's local Agent coding benchmark on RTX Pro 6000: 8B quantized model nearly matches 27B. Hardware math may need rewriting.

Aug 282 min read
RedditLocalLLaMA

Local LLMs Have a Hidden Bill Nobody Calculated — Your GPU Heats the Room

Reddit benchmarks: dual 5060Ti running Qwen 27B for coding hits 170W per card—equal to a small space heater running nonstop. Local AI looks "free," bu

Aug 282 min read