1 article tagged with this topic
A Reddit LocalLLaMA thread: Qwen 0.8B compresses conversation history, hands off to a 70B model for inference. Task-matched routing is quietly reshapi