找到 1 篇关于此标签的文章
Reddit LocalLLaMA 社区测试发现,Qwen 27B 在 12 万 token 长文本下,f16 精度 KV 缓存比 q8_0 更稳。量化不是省显存那么简单。