Alibaba
30 articles tagged with this topic
Qwen 1-bit is still 6x slower — companies eyeing local LLMs should wait
Qwen's 1-bit quantization runs 6x slower than the 4-bit version with ~70% accuracy. Local LLM deployment isn't ready to replace cloud APIs.
Engineer Pushes Qwen to the Limit: 260K Tokens Is Local AI's Hard Ceiling
Engineer pushed Qwen to extremes: context over 100K tokens drops generation 75%. Long-context remains local AI's hard ceiling—proof enterprises can't
Qwen 27B hits 50 tok/s: 16GB consumer GPUs can now run large models locally
A Reddit user ran Alibaba's Qwen 27B on a 16GB consumer GPU, hitting 50 tok/s generation + 100k context. Local LLMs are leaving the geek circle behind
Qwen Makes Thinking Depth Adjustable — Alibaba Lets LLMs Allocate Compute On Demand
Alibaba's Qwen now lets users adjust 'thinking depth'—quick answers for easy questions, more reasoning for hard ones. LLMs shift from on/off switch to
Alibaba Open-Sources Code Review Tool — The Real Win Isn't AI, It's Engineering
Alibaba open-sources OpenCodeReview, hitting 21k Stars. Hybrid architecture—engineering + LLM Agent—fixes three flaws of general AI agents in code rev
Alibaba Open-Source Model Revives a Bricked Foldable — Local AI Runs Solo
Reddit user revived a bricked foldable with Qwen 3.8 (27B) on a Raspberry Pi, saving ~$600. Real signal: local AI trusted solo on critical tasks.
Alibaba's Pixelle-Video: 5-Minute Demo Is Hype, the Pipeline Is the Playbook
Alibaba AIDC open-sources Pixelle-Video, a pluggable short-video pipeline. The 5-minute human time is the gimmick — the orchestration architecture is
Qwen 27B Runs on Just 18GB VRAM — Local LLM Bar Drops Again
Qwen's 27B multimodal needs only 18GB VRAM (Q4), runnable on consumer GPUs. Local LLM bar drops again, but MoE version still demands clusters.
A 24GB workstation card runs Qwen3 27B — local LLMs are finally viable
Reddit dev ran Qwen3 27B on a single 24GB workstation GPU, hitting 128K context and 60 tok/s. ~$2.8K hardware now handles mid-size LLMs locally.
Multimodal LLMs in Three Generations: China Leads Gen 2, Google Jumps to Gen 3
Multimodal LLMs crossed three architecture generations in three years. China leads Gen 2; Google's Gen 3 may be the next watershed.
Qwen Runs Locally — The "AI-Must-Be-Cloud" Assumption Cracks
Alibaba's Qwen3.8-Flash-Next lands in llama.cpp, letting regular PCs run AI locally without cloud APIs—open-source's quiet win over closed vendors.
Alibaba's Qwen Tops Both Open-Source Frontiers — Parameters Weren't the Question
Alibaba's Qwen hits the Pareto frontier on both total and active parameters. Running AI may keep getting cheaper — but the lead may not hold.
Reddit Pushes Alibaba 35B Re-Quantization — Open-Source Compute Drops a Notch
A Reddit request urges re-quantizing Alibaba's Qwen 35B with Unsloth's new UD 3.0 — open-source AI compute keeps getting cheaper.
Qwen 27B Compressed to 3-bit Runs 3D Apps Locally — Open Source Closes Cloud Gap
Reddit LocalLLaMA compressed Alibaba's Qwen 27B with 3-bit IQ3XXS, running a 3D demo locally. Small-model + quantization + local trends accelerate.
Qwen 27B Runs 3D Coding on Old GPU: Local AI Good Enough, Not Stable Yet
A Reddit dev ran Qwen 3.8 27B on a $200 used RTX 3090 and got working Three.js code. We see small models catching up to cloud—but not production-stabl
Qwen Flash Drops, Unsloth Adopts Day Zero — Open-Source AI's Pace Just Changed
Qwen Flash launched; Unsloth adopted it same day. The release-to-consumer-hardware gap shrank to under 24 hours. Enterprise IT should take note.
Qwen 27B Fixed a Real Git Repo — But Its Eval Method Is the Real Story
A dev tested Qwen 27B on a real Git repo. It fixed code, ran tests, recovered. The real story: the eval method—and 6 criteria China AI lacks.
Qwen 27B Quantization Test: Only 0.2% Accuracy Loss on a Single GPU
A Qwen 27B compression test on a single RTX 6000 found Q6 accuracy just 0.2% below Q8 while saving ~4GB VRAM. The local-LLM barrier is falling fast.
Qwen 27B Takes 4 Hours Locally, Cloud Claude 21 Minutes — The Gap Isn't Hardware
39K lines of C: local Qwen 3 27B in 4 hours, cloud Claude Opus 5 in 21 minutes. Same weights, two harnesses, both broken. Problem isn't the GPU.
Local Qwen 27B for Systems Programming? A Developer Pours Cold Water
Reddit developer asks: can local LLMs really write Rust/C++ system code, or only demo-friendly tasks? The old debate on overhyped AI coding.
Qwen3-27B on Single RTX 5090: Local LLMs Cross the Consumer Threshold
Developer runs quantized Qwen3-27B on a single RTX 5090: 450K context, vision preserved, 120 tok/s, 400W. Mid-tier LLMs no longer need multi-GPU serve
Qwen Crams a 9B Model into a 16GB Phone — China's First Real Say in On-Device AI
Reddit's LocalLLaMA names Alibaba's Qwen 9B as the top model for 16GB phones. On-device AI hits a tipping point — Chinese models earn their first real
Qwen Quantization Quality Varies 4×—Size Alone Isn’t Enough
A Reddit user tested 24 Qwen3.8-27B builds and found up to a 3–4× quality gap at 4-bit. File size alone is not a reliable guide.
Alibaba's Qwen Frustrates Local AI Users — They Want a Runnable Version
A Reddit user praises Qwen 3.8 27B quality but can't run it on M1 Max. They want 35B A3B — slower but usable. Local LLMs hit a hardware wall.
16GB VRAM Hits Flagship Performance — Closed-Source Giants Go Quiet
Qwen's new 27B open-source model runs locally on a single 16GB gaming GPU at near-flagship quality. Notably, OpenAI and Anthropic stayed silent this t
Qwen 3.8-27B One Week In: Top Marks for Doing, Memory Slips
Qwen 3.8-27B earned 'local best' on agent tasks in 2,000 Reddit tests but regressed on knowledge memory—first open-source model approaching GPT on exe
Alibaba packs Opus-class AI into 24GB GPUs; open-source local LLMs truly work
Qwen3.8-27B hit Hugging Face trending #1 in 48 hours; Cline made it default in 4 days. Open-source LLMs hit a usable threshold for the first time.
Qwen 27B Local Coding Test: Two AI Coding Agents, Surprisingly Wide Gap
On an RTX 3090, a developer ran two AI coding Agents on Qwen 27B. PI Agent won on resources and context, showing open-source LLMs can run Agent tasks.
Qwen Runs Faster and Cooler on 3090 — Local LLMs Are Finally Real Tools
Reddit user runs Qwen 27B on two RTX 3090s at 143 tokens/sec, dropping temps from 70°C to 35°C. Local LLMs cross from hobby to usable tool.
Alibaba Qwen Ships Twice in 6 Months — Local-Usable Builds Always Lag 2 Months
Alibaba's Qwen shipped two versions in six months, but the locally deployable versions always trail the benchmark leader by two months.