This week, OpenRouter and OpenCode launched a model called 'Space Bunny Alpha,' and Reddit user /u/crusaderky noticed something odd: its reasoning process (Chain of Thought—the intermediate step-by-step thinking) uses minimalist 'caveman language'—'the cat sat on the mat' gets compressed into 'cat sit mat.' This developer specifically uninstalled all plugins to confirm: this isn't a client-side issue, it's the model's own design. The model is widely believed to be MiniMax's M3.1. Our take: this isn't a bug, it's a cost-cutting strategy.

What this is

Last year, OpenAI o1 turned 'externalized thinking' into LLMs' new selling point—AI writes out its reasoning first, then delivers the final answer. Better results, but token consumption (models bill per input/output word) doubles. So-called 'caveman mode' compresses the language of that thinking process too: express the same reasoning with the fewest words. Space Bunny Alpha isn't an isolated case—in recent months, multiple open-weight (downloadable to run locally) reasoning models have been testing this approach.

Industry view

Supporters see this as commercialization's logical end state. Top vendors including Anthropic and DeepSeek are all racing to 'make reasoning cheaper'—because enterprise customers ultimately look at per-call cost, not how elegant the thinking process reads. Save tokens, push API prices down, and 'deep thinking' features can scale.

But researchers warn of three risks: first, over-compression may cause models to drop reasoning detail, increasing cases of 'right reasoning, wrong answer'; second, debugging gets harder—developers can't see what the model actually thought, so problems can only be guessed at; third, strong benchmark scores don't equal real-world stability, because benchmarks don't evaluate whether the thinking process itself is sound.

Impact on regular people

  • For enterprise IT: Over the next six months, per-call API pricing for reasoning models will likely keep falling—total cost for deep analysis and intelligent customer service will become more controllable.
  • For working professionals: When using AI for long-document analysis and code review, wait times will shrink and tasks completed per unit time will rise.
  • For consumer markets: Consumer products like ChatGPT and Claude may follow suit with compression strategies, letting 'deep thinking' features cover more users without price hikes.