返回首页

对比阅读

对比阅读:GPT-OSS Chat Template Bug: 20B Drifts in Long Chats, 120B Self-Recovers 与 GPT-OSS 模板藏 bug:开源大模型 20B 长对话跑偏,120B 能救回来

AEN
GPT-OSSOpenAIUnsloth·

GPT-OSS Chat Template Bug: 20B Drifts in Long Chats, 120B Self-Recovers

This week, developer arbv published a fixed chat template for GPT-OSS on Hugging Face, patching a hidden bug that causes the model to seriously drift in multi-turn conversations. We note: whether the open-source LLM experience feels good or bad increasingly hinges not on the model itself, but on the entire toolchain quietly determining success or failure.

What this is

GPT-OSS is OpenAI's open-source large language model released last year, in 20B and 120B variants (B stands for billion; parameters can be loosely understood as the model's "brain capacity"). The bug discovered this time lives in its chat template (a format instruction telling the model how to read conversation history): when the conversation history includes messages containing both "thinking" and "final answer" sections, the template only renders the thinking portion and silently drops the final answer content.

The consequence: the model sees incomplete content in earlier turns. A small model like 20B can't bear the load and drifts further the longer the conversation runs; while 120B, thanks to its larger size, can recover the conversational direction from the thinking traces, so the problem is less pronounced. The author also added a preserve_thinking option in the patch that actively preserves the thinking process, letting prefix caching (previously computed content in earlier turns doesn't need to be recomputed) run faster in multi-turn reasoning.

Industry view

The community broadly welcomes the fix, and many believe it neatly explains "why so many people trying out GPT-OSS 20B felt it was clearly worse than the official marketing." A new consensus is forming: the open-source model moat has shifted from "can the weights be downloaded" to "is the toolchain stable." Others push back — this bug is too deep and too specialized for normal users to even reach that turn; meanwhile the author himself admits that OpenAI's official reference template doesn't have this issue, the blame mainly falls on the third-party derivative Unsloth. The takeaway for everyone: when running community-modified versions, a little extra vigilance never hurts.

Impact on regular people

  • For enterprise IT: open-source models aren't usable just by downloading weights; you also need to invest engineering hours in templates, inference, and alignment. Budgets can't just count GPU costs.
  • For individual professionals: when building workflows on local or open-source LLMs, encountering "the model suddenly going dumb" may not be the model's fault — it could be a template or front-end toolchain pitfall.
  • For the consumer market: the watershed for open-source LLMs is shifting from "can it run" to "is long-conversation stable." Going forward, vendors will sell not just parameters, but stable and reliable toolchains.
BZH
GPT-OSSOpenAIUnsloth·

GPT-OSS 模板藏 bug:开源大模型 20B 长对话跑偏,120B 能救回来

本周,开发者 arbv 在 Hugging Face 上发布了一个 GPT-OSS 的修正聊天模板,修的是一个让模型在多轮对话中严重跑偏的隐藏 bug。我们注意到:开源大模型的体验好不好,越来越不只是模型本身的事,整个工具链都在悄悄决定成败。

这是什么

GPT-OSS 是 OpenAI 去年开源的大语言模型,分 20B 和 120B 两个版本(B 代表 billion,即 10 亿参数,参数可粗略理解为模型的"脑容量")。这次被发现的是它聊天模板(一种告诉模型如何阅读对话历史的格式说明)里的 bug:当对话历史里有同时包含「思考」和「最终回答」的消息时,模板只渲染了思考部分,把最终回答内容丢掉了。

后果是:模型在前几轮看到的内容是残缺的。20B 这种小模型扛不住,越聊越偏;而 120B 因为体量大,能从思考痕迹里自己找回对话方向,所以问题不那么明显。作者也在补丁里加了 preserve_thinking 选项,主动保留思考过程,让多轮推理的前缀缓存(前几轮算过的内容不用重算)能跑得更快。

行业怎么看

社区普遍欢迎这个修复,不少人认为它正好解释了"为什么很多人试用 GPT-OSS 20B 觉得明显不如官方宣传"。一种新的判断正在形成:开源模型的护城河已经从"权重能不能下载"转向"工具链稳不稳"。也有人提出反对看法——这个 bug 太深、太专业,普通用户根本聊不到那个轮次;同时作者自己也承认,OpenAI 官方参考模板并没有这个问题,锅主要在第三方衍生版 Unsloth 上。这给所有人的提醒是:用社区魔改版本时,多一份警觉总没坏处。

对普通人的影响

  • 对企业 IT:开源模型不是下载权重就能用,还得投入人力做模板、推理、对齐的二次开发,预算不能只算显卡钱。
  • 对个人职场:用本地或开源大模型搭工作流时,遇到"模型突然变笨"未必是模型本身的锅,可能是模板或前端工具链的坑。
  • 对消费市场:开源大模型的分水岭正从"能不能跑"转向"长对话稳不稳",未来厂商卖的不仅是参数,更是稳定可靠的工具链。