1 article tagged with this topic
Reddit LocalLLaMA tests show Qwen 27B with f16 KV cache outperforms q8_0 at 120K-token contexts. Quantization isn't just about saving VRAM.