Reasoning Models
5 articles tagged with this topic
Qwen 3.8 Tested: Deep Thinking Burns 5.5x Tokens—Local Deployment Math Changes
Reddit user tested Qwen3.8-27B on M5 Max: deep thinking uses 5.5x tokens, 6x time; disabling tanks quality. The "thinking" cost gap is exposed.
IBM Releases Three Open-Source Reasoning Models, Free for Commercial Use
IBM drops Granite 4.2 reasoning models (30B/8B/3B) on Hugging Face under Apache 2.0, free for commercial use. Supports chain-of-thought and 512K conte
Qwen Users Roast Reasoning Models: 50% of 'Thinking' Is Just 'Wait' Tokens
Reddit joke exposes a real problem: reasoning models' thinking chains are filled with filler like 'wait', bloating KV cache and exploding deployment c
RL Changes Just 1-3% of Output — Reasoning Model Training May Be 1000x Overpriced
Viral paper: RL only changes 1-3% of reasoning model output. Drop RL, and you may get similar reasoning at 1/1000 the compute cost.
Reddit User Hacks 'Lite Mode' into Open-Source 27B Model, Cuts Compute 5x
Reddit user crafts new tier for 27B reasoning model — mixes 'low' and 'max' prompts, quality near max but thinking tokens at 1/5. Update for local AI