Back to home
千问
2 articles tagged with this topic
Qwen千问
Context Is Burning Cash — Why 0.8B Models Are Easing 70B's Load
A Reddit LocalLLaMA thread: Qwen 0.8B compresses conversation history, hands off to a 70B model for inference. Task-matched routing is quietly reshapi
Aug 232 min read
Qwen千问
Chinese Open-Source LLMs Are Now Slimmable: Qwen 27B Trimmed to 23B, No Retraining
A Reddit dev shrank Qwen 27B to ~23B by stripping layers — no retraining. Open-source Chinese LLMs are now "hackable" for local deployment.
Aug 202 min read