What this is

This week on Reddit's open-source model community LocalLLaMA, a set of comparison tests sparked heated debate: the latest-generation model (3.8) from the same lab scored worse than its predecessor (3.6) when "thinking mode" — the feature that lets the AI write a long reasoning passage before answering — was disabled during benchmarking.

The poster GodComplecs judged that, in pursuit of betting on "agents" (Agent — the form of AI that plans its own steps and completes tasks autonomously), labs are sacrificing the "base model" experience. Base model refers to versions that respond directly to instructions without relying on long reasoning — the form most commonly used in enterprises for writing emails, editing copy, and producing summaries.

Industry view

The mainstream direction remains betting on thinking mode. OpenAI's o-series, Anthropic Claude, DeepSeek R1, and Alibaba Qwen are all making "think first, then answer" the default. The underlying logic: teaching AI to decompose complex tasks is the path to higher intelligence.

Opposition clusters around three points. First, thinking has costs — calls are slower, tokens more expensive, wasteful for email-writing and copy-editing scenarios. Second, community tests show thinking models perform worse than pure instruction models with thinking turned off, indicating base capabilities are being squeezed. Third, Reddit commenters questioned whether labs are concentrating compute on the "sexier" agent narrative while neglecting the features ordinary users rely on most.

In our view, the risk is that as every top lab does this, "lightweight, usable, cheap" non-thinking models in the open-source ecosystem may dwindle — and the ones paying the bill will be enterprise IT budgets and end-user experience.

Impact on regular people

Enterprise IT procurement: when selecting models, distinguish between "thinking mode" and "normal mode" and match by scenario. Newer isn't always better.

Workplace productivity: for daily tasks like writing emails, editing copy, and producing summaries, long thinking drags out response times. Consider actively turning the feature off.

Consumer market: consumer-grade AI products (smart-assistant categories) may become slower and pricier. Because "thinking" consumes more compute, those costs ultimately pass through to subscription prices.