Back to home

Alibaba

30 articles tagged with this topic

QwenAlibaba

Qwen 1-bit is still 6x slower — companies eyeing local LLMs should wait

Qwen's 1-bit quantization runs 6x slower than the 4-bit version with ~70% accuracy. Local LLM deployment isn't ready to replace cloud APIs.

1h ago2 min read
QwenAlibaba

Engineer Pushes Qwen to the Limit: 260K Tokens Is Local AI's Hard Ceiling

Engineer pushed Qwen to extremes: context over 100K tokens drops generation 75%. Long-context remains local AI's hard ceiling—proof enterprises can't

3h ago2 min read
QwenTongyi Qianwen

Qwen 27B hits 50 tok/s: 16GB consumer GPUs can now run large models locally

A Reddit user ran Alibaba's Qwen 27B on a 16GB consumer GPU, hitting 50 tok/s generation + 100k context. Local LLMs are leaving the geek circle behind

13h ago2 min read
QwenAlibaba

Qwen Makes Thinking Depth Adjustable — Alibaba Lets LLMs Allocate Compute On Demand

Alibaba's Qwen now lets users adjust 'thinking depth'—quick answers for easy questions, more reasoning for hard ones. LLMs shift from on/off switch to

17h ago2 min read
AlibabaOpenCodeReview

Alibaba Open-Sources Code Review Tool — The Real Win Isn't AI, It's Engineering

Alibaba open-sources OpenCodeReview, hitting 21k Stars. Hybrid architecture—engineering + LLM Agent—fixes three flaws of general AI agents in code rev

17h ago2 min read
QwenAlibaba

Alibaba Open-Source Model Revives a Bricked Foldable — Local AI Runs Solo

Reddit user revived a bricked foldable with Qwen 3.8 (27B) on a Raspberry Pi, saving ~$600. Real signal: local AI trusted solo on critical tasks.

21h ago2 min read
AlibabaPixelle-Video

Alibaba's Pixelle-Video: 5-Minute Demo Is Hype, the Pipeline Is the Playbook

Alibaba AIDC open-sources Pixelle-Video, a pluggable short-video pipeline. The 5-minute human time is the gimmick — the orchestration architecture is

21h ago2 min read
QwenAlibaba

Qwen 27B Runs on Just 18GB VRAM — Local LLM Bar Drops Again

Qwen's 27B multimodal needs only 18GB VRAM (Q4), runnable on consumer GPUs. Local LLM bar drops again, but MoE version still demands clusters.

1d ago2 min read
Qwen3Alibaba

A 24GB workstation card runs Qwen3 27B — local LLMs are finally viable

Reddit dev ran Qwen3 27B on a single 24GB workstation GPU, hitting 128K context and 60 tok/s. ~$2.8K hardware now handles mid-size LLMs locally.

1d ago2 min read
Qwen-VLGemma4

Multimodal LLMs in Three Generations: China Leads Gen 2, Google Jumps to Gen 3

Multimodal LLMs crossed three architecture generations in three years. China leads Gen 2; Google's Gen 3 may be the next watershed.

2d ago2 min read
QwenAlibaba

Qwen Runs Locally — The "AI-Must-Be-Cloud" Assumption Cracks

Alibaba's Qwen3.8-Flash-Next lands in llama.cpp, letting regular PCs run AI locally without cloud APIs—open-source's quiet win over closed vendors.

2d ago2 min read
QwenAlibaba

Alibaba's Qwen Tops Both Open-Source Frontiers — Parameters Weren't the Question

Alibaba's Qwen hits the Pareto frontier on both total and active parameters. Running AI may keep getting cheaper — but the lead may not hold.

2d ago2 min read
QwenUnsloth

Reddit Pushes Alibaba 35B Re-Quantization — Open-Source Compute Drops a Notch

A Reddit request urges re-quantizing Alibaba's Qwen 35B with Unsloth's new UD 3.0 — open-source AI compute keeps getting cheaper.

2d ago2 min read
QwenTongyi Qianwen

Qwen 27B Compressed to 3-bit Runs 3D Apps Locally — Open Source Closes Cloud Gap

Reddit LocalLLaMA compressed Alibaba's Qwen 27B with 3-bit IQ3XXS, running a 3D demo locally. Small-model + quantization + local trends accelerate.

3d ago2 min read
QwenAlibaba

Qwen 27B Runs 3D Coding on Old GPU: Local AI Good Enough, Not Stable Yet

A Reddit dev ran Qwen 3.8 27B on a $200 used RTX 3090 and got working Three.js code. We see small models catching up to cloud—but not production-stabl

4d ago2 min read
QwenUnsloth

Qwen Flash Drops, Unsloth Adopts Day Zero — Open-Source AI's Pace Just Changed

Qwen Flash launched; Unsloth adopted it same day. The release-to-consumer-hardware gap shrank to under 24 hours. Enterprise IT should take note.

4d ago2 min read
QwenTongyi Qianwen

Qwen 27B Fixed a Real Git Repo — But Its Eval Method Is the Real Story

A dev tested Qwen 27B on a real Git repo. It fixed code, ran tests, recovered. The real story: the eval method—and 6 criteria China AI lacks.

5d ago2 min read
QwenTongyi Qianwen

Qwen 27B Quantization Test: Only 0.2% Accuracy Loss on a Single GPU

A Qwen 27B compression test on a single RTX 6000 found Q6 accuracy just 0.2% below Q8 while saving ~4GB VRAM. The local-LLM barrier is falling fast.

6d ago2 min read
QwenClaude

Qwen 27B Takes 4 Hours Locally, Cloud Claude 21 Minutes — The Gap Isn't Hardware

39K lines of C: local Qwen 3 27B in 4 hours, cloud Claude Opus 5 in 21 minutes. Same weights, two harnesses, both broken. Problem isn't the GPU.

6d ago2 min read
QwenAlibaba

Local Qwen 27B for Systems Programming? A Developer Pours Cold Water

Reddit developer asks: can local LLMs really write Rust/C++ system code, or only demo-friendly tasks? The old debate on overhyped AI coding.

6d ago2 min read
AlibabaQwen

Qwen3-27B on Single RTX 5090: Local LLMs Cross the Consumer Threshold

Developer runs quantized Qwen3-27B on a single RTX 5090: 450K context, vision preserved, 120 tok/s, 400W. Mid-tier LLMs no longer need multi-GPU serve

6d ago2 min read
QwenAlibaba

Qwen Crams a 9B Model into a 16GB Phone — China's First Real Say in On-Device AI

Reddit's LocalLLaMA names Alibaba's Qwen 9B as the top model for 16GB phones. On-device AI hits a tipping point — Chinese models earn their first real

6d ago2 min read
QwenTongyi Qianwen

Qwen Quantization Quality Varies 4×—Size Alone Isn’t Enough

A Reddit user tested 24 Qwen3.8-27B builds and found up to a 3–4× quality gap at 4-bit. File size alone is not a reliable guide.

6d ago2 min read
AlibabaQwen

Alibaba's Qwen Frustrates Local AI Users — They Want a Runnable Version

A Reddit user praises Qwen 3.8 27B quality but can't run it on M1 Max. They want 35B A3B — slower but usable. Local LLMs hit a hardware wall.

6d ago2 min read
Qwenopen-source model

16GB VRAM Hits Flagship Performance — Closed-Source Giants Go Quiet

Qwen's new 27B open-source model runs locally on a single 16GB gaming GPU at near-flagship quality. Notably, OpenAI and Anthropic stayed silent this t

Aug 232 min read
QwenAlibaba

Qwen 3.8-27B One Week In: Top Marks for Doing, Memory Slips

Qwen 3.8-27B earned 'local best' on agent tasks in 2,000 Reddit tests but regressed on knowledge memory—first open-source model approaching GPT on exe

Aug 232 min read
AlibabaQwen3.8

Alibaba packs Opus-class AI into 24GB GPUs; open-source local LLMs truly work

Qwen3.8-27B hit Hugging Face trending #1 in 48 hours; Cline made it default in 4 days. Open-source LLMs hit a usable threshold for the first time.

Aug 222 min read
QwenAlibaba

Qwen 27B Local Coding Test: Two AI Coding Agents, Surprisingly Wide Gap

On an RTX 3090, a developer ran two AI coding Agents on Qwen 27B. PI Agent won on resources and context, showing open-source LLMs can run Agent tasks.

Aug 222 min read
QwenvLLM

Qwen Runs Faster and Cooler on 3090 — Local LLMs Are Finally Real Tools

Reddit user runs Qwen 27B on two RTX 3090s at 143 tokens/sec, dropping temps from 70°C to 35°C. Local LLMs cross from hobby to usable tool.

Aug 222 min read
AlibabaQwen

Alibaba Qwen Ships Twice in 6 Months — Local-Usable Builds Always Lag 2 Months

Alibaba's Qwen shipped two versions in six months, but the locally deployable versions always trail the benchmark leader by two months.

Aug 222 min read