A coding benchmark on Reddit this week targeting Qwen3.6-35B shows that only 1 of 5 community fine-tunes matched the original — this "game-laptop-runnable" open-source model from Alibaba's Tongyi Qianwen has fallen into an embarrassing community-wide "worse with each fix" pattern.
What this is
Qwen3.6-35B-A3B is a "small but mighty" model released by Alibaba's Tongyi Qianwen: 35 billion total parameters, but only about 3 billion activated at inference. Technically, it's a MoE (Mixture of Experts — an architecture that lets the model "split into teams" internally to save compute). VRAM usage sits near the 3-billion tier, so it runs on an ordinary gaming laptop.
The tester ran Aider Polyglot, a coding benchmark, head-to-head against the original and 5 community fine-tunes. The results were counterintuitive: the original scored 37.4% on first pass and 71% with retries. Only Occamy-1.0 came close among the fine-tunes — the rest all trailed. Tiel, which had a solid community reputation, performed worst, at just 53.3% pass-after-retry.
Industry view
Supporters argue that small-scale fine-tuning isn't always additive — it can damage the base model's general capabilities. The vendor's pretraining data scale and training investment vastly exceed what any community can mobilize, so blindly chasing "optimized for one scenario" often backfires.
But dissent exists. Critics point out that the benchmark only covers coding — it says nothing about writing, reasoning, or other tasks. A sample size of 5 is also too small to support a sweeping "community fine-tunes all regress" conclusion. Other developers note that prompt templates (the format in which the model receives instructions) significantly affect scores, and swapping the template could change the results — meaning the test's stability is questionable.
Impact on regular people
For enterprise IT: deploying open-source models locally is becoming viable, but "original vs. fine-tuned" is now a real choice. Procurement and engineering teams need to re-evaluate their selection strategies.
For working professionals: if you plan to use a local model for daily work, the most pragmatic path right now is to get the workflow running on the original Qwen3.6-35B first, then decide whether customization for specific tasks is worth the effort.
For the consumer market: the very fact that open-source models run on small devices is shrinking the necessity of "must call cloud APIs." The cost structure of personal AI tools could shift going forward.