Back to home
long context
2 articles tagged with this topic
KV cacheLocalLLaMA
Local Hack Triples 256k Inference Speed, but Commercial Hurdles Remain
Reddit user validates chunked KV cache on a small open-source model: 256k prefill runs 3x faster, needle-in-haystack accuracy holds. Real value: bottl
Aug 222 min read
DeepSeekV4-Flash
DeepSeek V4-Flash Falls Asleep Mid-Task: The Hidden Cost of Local AI Deployment
DeepSeek's V4-Flash open-source model silently halts generation past 100k tokens. Concrete proof that "downloadable" and "reliably usable" are still f
Aug 92 min read