This week, developer arbv published a fixed chat template for GPT-OSS on Hugging Face, patching a hidden bug that causes the model to seriously drift in multi-turn conversations. We note: whether the open-source LLM experience feels good or bad increasingly hinges not on the model itself, but on the entire toolchain quietly determining success or failure.

What this is

GPT-OSS is OpenAI's open-source large language model released last year, in 20B and 120B variants (B stands for billion; parameters can be loosely understood as the model's "brain capacity"). The bug discovered this time lives in its chat template (a format instruction telling the model how to read conversation history): when the conversation history includes messages containing both "thinking" and "final answer" sections, the template only renders the thinking portion and silently drops the final answer content.

The consequence: the model sees incomplete content in earlier turns. A small model like 20B can't bear the load and drifts further the longer the conversation runs; while 120B, thanks to its larger size, can recover the conversational direction from the thinking traces, so the problem is less pronounced. The author also added a preserve_thinking option in the patch that actively preserves the thinking process, letting prefix caching (previously computed content in earlier turns doesn't need to be recomputed) run faster in multi-turn reasoning.

Industry view

The community broadly welcomes the fix, and many believe it neatly explains "why so many people trying out GPT-OSS 20B felt it was clearly worse than the official marketing." A new consensus is forming: the open-source model moat has shifted from "can the weights be downloaded" to "is the toolchain stable." Others push back — this bug is too deep and too specialized for normal users to even reach that turn; meanwhile the author himself admits that OpenAI's official reference template doesn't have this issue, the blame mainly falls on the third-party derivative Unsloth. The takeaway for everyone: when running community-modified versions, a little extra vigilance never hurts.

Impact on regular people

  • For enterprise IT: open-source models aren't usable just by downloading weights; you also need to invest engineering hours in templates, inference, and alignment. Budgets can't just count GPU costs.
  • For individual professionals: when building workflows on local or open-source LLMs, encountering "the model suddenly going dumb" may not be the model's fault — it could be a template or front-end toolchain pitfall.
  • For the consumer market: the watershed for open-source LLMs is shifting from "can it run" to "is long-conversation stable." Going forward, vendors will sell not just parameters, but stable and reliable toolchains.