返回首页

对比阅读

对比阅读:Qwen3.6-35B Fine-tunes Trail the Original — The 'Worse with Each Fix' Trap 与 Qwen3.6-35B 微调版多数不如原版 — 社区'越改越差'困局浮现

AEN
QwenTongyi QianwenAlibaba·

Qwen3.6-35B Fine-tunes Trail the Original — The 'Worse with Each Fix' Trap

A coding benchmark on Reddit this week targeting Qwen3.6-35B shows that only 1 of 5 community fine-tunes matched the original — this "game-laptop-runnable" open-source model from Alibaba's Tongyi Qianwen has fallen into an embarrassing community-wide "worse with each fix" pattern.

What this is

Qwen3.6-35B-A3B is a "small but mighty" model released by Alibaba's Tongyi Qianwen: 35 billion total parameters, but only about 3 billion activated at inference. Technically, it's a MoE (Mixture of Experts — an architecture that lets the model "split into teams" internally to save compute). VRAM usage sits near the 3-billion tier, so it runs on an ordinary gaming laptop.

The tester ran Aider Polyglot, a coding benchmark, head-to-head against the original and 5 community fine-tunes. The results were counterintuitive: the original scored 37.4% on first pass and 71% with retries. Only Occamy-1.0 came close among the fine-tunes — the rest all trailed. Tiel, which had a solid community reputation, performed worst, at just 53.3% pass-after-retry.

Industry view

Supporters argue that small-scale fine-tuning isn't always additive — it can damage the base model's general capabilities. The vendor's pretraining data scale and training investment vastly exceed what any community can mobilize, so blindly chasing "optimized for one scenario" often backfires.

But dissent exists. Critics point out that the benchmark only covers coding — it says nothing about writing, reasoning, or other tasks. A sample size of 5 is also too small to support a sweeping "community fine-tunes all regress" conclusion. Other developers note that prompt templates (the format in which the model receives instructions) significantly affect scores, and swapping the template could change the results — meaning the test's stability is questionable.

Impact on regular people

For enterprise IT: deploying open-source models locally is becoming viable, but "original vs. fine-tuned" is now a real choice. Procurement and engineering teams need to re-evaluate their selection strategies.

For working professionals: if you plan to use a local model for daily work, the most pragmatic path right now is to get the workflow running on the original Qwen3.6-35B first, then decide whether customization for specific tasks is worth the effort.

For the consumer market: the very fact that open-source models run on small devices is shrinking the necessity of "must call cloud APIs." The cost structure of personal AI tools could shift going forward.

BZH
Qwen通义千问阿里巴巴·

Qwen3.6-35B 微调版多数不如原版 — 社区'越改越差'困局浮现

本周 Reddit 上一项针对 Qwen3.6-35B 的编程实测显示:5 个社区微调版本中只有 1 个能与原版持平——这款阿里通义千问推出的"游戏本能跑"的开源模型,正陷入社区"越改越差"的尴尬。

这是什么

Qwen3.6-35B-A3B 是阿里通义千问上线的"小钢炮"级模型:总参数 350 亿,但推理时只激活约 30 亿——技术上属于 MoE(混合专家架构,让模型内部"分组答题"以节省算力),显存占用接近 30 亿级别,普通游戏本即可运行。

测试者用 Aider Polyglot 编程评测工具对它和 5 个社区微调版本做横向对比。结果反直觉:原版一次性通过率 37.4%、重试后达 71%;5 个微调版里只有 Occamy-1.0 接近这个水平,其余全部落后。社区口碑不错的 Tiel 反而最差,重试通过率仅 53.3%。

行业怎么看

支持者的判断是:小规模微调并不总是"加法",反而可能损害原模型的通用能力。原厂的预训练数据规模和训练投入远超社区能调动的资源,盲目追求"为某个场景优化"往往得不偿失。

但反对声音同样存在。有人指出,本次评测只覆盖编程一项任务,无法推及写作、推理等场景;5 个样本量也偏小,难以得出"社区微调整体退步"的结论。还有开发者认为,提示词模板(即模型接收指令的格式)对成绩影响显著,换一套模板结果可能不同——这意味着测试结论的稳定性存疑。

对普通人的影响

对企业 IT:本地部署开源模型正变得可行,但"用原版还是微调版"已是真实选择题,采购和工程团队需要重新评估选型策略。

对个人职场:若打算用本地模型处理日常工作,目前最务实的路径是先用原版 Qwen3.6-35B 跑通流程,再判断是否值得针对特定任务做定制。

对消费市场:开源模型在小设备上跑得动这件事本身,正在压缩"必须联网调用云端 API"的必要性,未来个人 AI 工具的成本结构可能改变。