A community developer this week tested on a Qwen 4B model: after disabling hesitation words like "wait" and "maybe", MATH-500 accuracy on a commonly used precision rose from 60% to 66%, while reasoning tokens (the model's intermediate thinking text) actually dropped 14.8%. This is no folk remedy — we note that a Meta paper already pointed out that frequent use of these words during AI problem-solving is a signal of "overthinking", and blocking them yields better accuracy.
What this is
The specific approach uses llama.cpp's (a tool for locally running open-source large models) logit bias (a feature that adjusts output probabilities for specific words) to suppress nearly 50 hesitation words including "wait", "maybe", and "perhaps". The developer tested five precisions — BF16, Q8_0, Q4_K_M, Q3_K_M, Q2_K (smaller numbers mean heavier compression, lower VRAM usage but greater precision loss). Results were remarkably consistent: BF16 rose from 74% to 84%, Q2_K (the extreme compression version) doubled from 12% to 24%, with nearly all precisions showing 6–14 percentage point gains.
Plain language: stop AI from talking to itself while solving problems, and answers get more accurate while costs drop.
Industry view
Supporters are excited. This is a win for the open-source community — big companies publish papers, ordinary developers reproduce and extend them across different precisions within a week. Reasoning tokens down 10–20% means the same hardware can run more tasks, direct good news for enterprises sensitive to local deployment costs.
Opposition is also clear. First, this is a single-model, single-test-set (MATH-500) result; conclusions cannot be directly generalized to all tasks. Second, the method is crude — normal transition words like "however" and "but" also get blocked, potentially damaging language coherence. Third, in chat scenarios no one wants to talk to an AI that won't say "maybe"; the experience becomes "robotic".
Impact on regular people
For enterprise IT: Hidden costs of locally deploying open-source large models may continue to fall, lowering the bar again for small and medium companies building their own AI assistants.
For professionals: Try this trick when using open-source models like Qwen for data analysis or reports, but don't expect similar gains in copywriting or dialogue scenarios.
For consumer markets: Applications involving AI solving math problems and exam tutoring will become more reliable; chatbot "tone" may turn mechanical.