Qwen
30 articles tagged with this topic
Qwen 27B on a Pi Dumps Cloud APIs — Open-Source AI Is Finally 'Good Enough'
Dev runs Qwen 27B on Pi, abandons OpenAI and Anthropic APIs. Local cost matches cheapest cloud—a reset signal for open-source AI's commercial value.
Deploying LLMs Means Tuning 40+ Parameters—It's Becoming an Independent Craft
40+ parameters to tune per LLM deployment: a new guide reveals AI production has hit engineering hard reality—tuning is now a standalone craft.
Aliyun Expands to Turkey, Finland, Netherlands — Export Edge Isn't the Model
Aliyun adds data centers in Turkey, Finland, Netherlands. The real AI-export bottleneck isn't models—it's cloud geography, the compliance and cost var
Dewu Rebuilds Search With LLMs — But Generative Recall's Cost Is Unclear
Dewu's 'generative recall' rebuilds search with LLM-generated tags—the first systematic engineering writeup from a major Chinese e-commerce team.
$7K Mac Studio Hits 1,700 Agent Schedules — Local LLMs Finally Do Real Work
Reddit dev runs Qwen 3 on M5 Ultra Mac Studio (96GB), hits 1,700 sub-agent schedules and 89% cache hit on 112M-token dev task — first engineering-grad
40MB Single File Packs LLM + Voice — Local-First AI Goes Live
Indie dev packed LLM, voice, code execution into a 40MB local exe; AI generates every UI pixel live. Rare working Local-First demo challenging cloud v
Qwen 27B Now Runs on One AMD GPU — Local LLMs Start Feeling Installable
Reddit dev ships Qwen 27B GGUF quantization on a single AMD Strix Halo, beating ISTA and Unsloth on three benchmarks. One person, one GPU, one week.
RTX 4090 Runs Qwen Locally — But Local AI Isn't Ready for White-Collar Work
A Reddit user ran Qwen 27B on an RTX 4090 for coding. Chinese open-source gains traction with developers; local LLMs remain a hobbyist toy.
Overseas Devs Build Qwen Coding Tools — Alibaba's Open-Source Bet Pays Off
Reddit's r/LocalLLaMA debated which 'harness' fits Qwen 3.8 best. Surface-level tool fight; underneath, Qwen is now a Western open-source default.
Reddit User Deploys AI Agent at Home — Early Local LLM Experiments
A popular Reddit post shows a home AI Agent setup: open-webui + isolated VM + Qwen LLM. We examine what it reveals about local LLMs leaving data cente
Apple M5 Max hits 350K-token context — hardware bar cleared, quality still lags
128GB MacBook Pro holds Qwen3.8 at 350K tokens, but past 100K the model 'confuses roles.' Hardware bar falls; quality still lags.
Qwen 1-bit is still 6x slower — companies eyeing local LLMs should wait
Qwen's 1-bit quantization runs 6x slower than the 4-bit version with ~70% accuracy. Local LLM deployment isn't ready to replace cloud APIs.
Engineer Pushes Qwen to the Limit: 260K Tokens Is Local AI's Hard Ceiling
Engineer pushed Qwen to extremes: context over 100K tokens drops generation 75%. Long-context remains local AI's hard ceiling—proof enterprises can't
Local AI Coding Is Trending — But Most Companies' GPUs Can't Run It
A Reddit post about running Qwen 3 27B locally on an RTX A4500 for AI coding sparked debate. Local model coding is shifting from hobbyist toy to real
Million-Token Context on Two 5090s — Amateur Dev Shatters Enterprise AI Myth
Reddit developer NInfer hits 1.04M token context on consumer RTX 5090s at 119 tok/s — 2.8x faster than vLLM with a 27B Qwen model.
Qwen 3.8 Tested: Deep Thinking Burns 5.5x Tokens—Local Deployment Math Changes
Reddit user tested Qwen3.8-27B on M5 Max: deep thinking uses 5.5x tokens, 6x time; disabling tanks quality. The "thinking" cost gap is exposed.
RTX 5080 Hits Just 6 Characters/Second on Qwen — How High Is Local AI's Home Threshold?
Reddit user hit ~6 chars/sec running quantized Qwen on RTX 5080 + 64GB RAM. Local AI still demands far more hardware and tuning than ordinary users ca
Community Qwen3.8 Quantization Saves 30GB — Local LLM Bar Drops Again
Community dev agentionai ships a custom quantized Qwen3.8-Flash-Next, 20–30GB smaller than mainstream versions at comparable quality. Local LLM bar dr
Qwen 27B hits 50 tok/s: 16GB consumer GPUs can now run large models locally
A Reddit user ran Alibaba's Qwen 27B on a 16GB consumer GPU, hitting 50 tok/s generation + 100k context. Local LLMs are leaving the geek circle behind
Maxing AI 'Thinking Depth' Hurts Results — DGX Spark Local Test Warns Enterprises
A Reddit developer tested DeepSeek/Qwen on four DGX Sparks: 'deep thinking' mode lowers scores and doubles runtime — a direct cost warning for AI infe
Qwen Makes Thinking Depth Adjustable — Alibaba Lets LLMs Allocate Compute On Demand
Alibaba's Qwen now lets users adjust 'thinking depth'—quick answers for easy questions, more reasoning for hard ones. LLMs shift from on/off switch to
Alibaba and Zhipu Bet on Small Models — Local AI Faces Choice Overload
Qwen Flash and GLM Flash launched together, leaving local users with choice overload. China's open-source LLMs shift from parameter wars to same-tier
Alibaba Open-Source Model Revives a Bricked Foldable — Local AI Runs Solo
Reddit user revived a bricked foldable with Qwen 3.8 (27B) on a Raspberry Pi, saving ~$600. Real signal: local AI trusted solo on critical tasks.
Qwen 27B Runs on Just 18GB VRAM — Local LLM Bar Drops Again
Qwen's 27B multimodal needs only 18GB VRAM (Q4), runnable on consumer GPUs. Local LLM bar drops again, but MoE version still demands clusters.
Alibaba Qwen 27B Squeezed to 10GB, Matches Original Quality — Local AI Advances
Austria's ISTA-DASLab squeezed a Qwen 27B to 10GB, matching original quality. Local AI advances — capable, private deployment without cloud uploads.
Two DGX Sparks Hit 181 tok/s Concurrent: Local Multi-Agent Is Now Viable
Two NVIDIA DGX Sparks + Qwen models hit 181 tok/s concurrent across 9 parallel AI Agents. Local multi-Agent moves from demo to real work.
Zhipu GLM-5.3 Scores Higher Without New Architecture — Chinese LLMs Turn Inward
Zhipu's GLM-5.3 hits stronger benchmarks with architecture identical to 5.2 — a training-only upgrade worth watching while peers chase new designs.
Two 3090s Run 27B Model at 165 tok/s — Local AI Is Finally 'Good Enough'
We noted a Reddit user hit 165 tok/s on a 27B Qwen model using two RTX 3090s (used rig under ¥20K) — Agent-ready. Local LLMs just crossed from 'toy' t
Local 8B Model Ties 27B on Agent Coding — The 'Good Enough' Moment Arrives
Reddit dev's local Agent coding benchmark on RTX Pro 6000: 8B quantized model nearly matches 27B. Hardware math may need rewriting.
Local LLMs Have a Hidden Bill Nobody Calculated — Your GPU Heats the Room
Reddit benchmarks: dual 5060Ti running Qwen 27B for coding hits 170W per card—equal to a small space heater running nonstop. Local AI looks "free," bu