KV cacheLocalLLaMA
Local Hack Triples 256k Inference Speed, but Commercial Hurdles Remain
Reddit user validates chunked KV cache on a small open-source model: 256k prefill runs 3x faster, needle-in-haystack accuracy holds. Real value: bottl
Aug 22·2 min read