A thread on LocalLLaMA has been surging all week: between Alibaba's open-source Qwen 27B and 35B-A3B, which one actually wins at coding, web scraping, and tool use? The poster dug around and found zero public benchmarks. That matters more to us than any answer: local LLMs have slid from geek toys into production tools, yet selection still runs on guesswork.

What this is

Alibaba has been flooding the zone with small open-source models marketed as "runnable on ordinary GPUs." The two versions differ in both parameter count and quantization (quantization = a trick that shrinks model file size; smaller files run faster but lose some quality). The 27B uses IQ3_XXS (extreme compression); the 35B-A3B uses Q4_K_M (light compression). In theory, the harder you compress, the faster it runs — but nobody has publicly measured how much quality actually drops on coding and data-scraping tasks, the ones that actually "do work."

Industry view

The local LLM community has split into two camps. One insists bigger parameters mean better results — 35B beats 27B, no contest. The other argues quantization matters more than parameter count, and that aggressive IQ3 compression routinely breaks down on coding tasks. But the dissenting voice worth listening to cuts deeper: raw model size tells you nothing. What actually determines whether a model can "do work" is the underlying training data distribution — and Alibaba hasn't disclosed that. Our read: users are picking models blind.

Impact on regular people

For SMB IT teams: you can now genuinely run an AI that codes and scrapes on your own servers — but without credible benchmarks at the selection stage, trial-and-error costs aren't low. Teams with budgets under ¥100,000 should pause before jumping in.

For individual professionals: non-technical readers can skip this, but know one thing — if someone at your company is tinkering with "local AI," they're almost certainly wrestling with exactly this kind of model-picking dilemma.

For consumer markets: no impact yet. The AI assistants on your phone are still cloud-based LLMs. Consumer-grade on-device AI is still two to three years out.