Back to home
local-LLM
2 articles tagged with this topic
DeepSeekspeculative-decoding
Speculative decoding is becoming standard — open-source LLMs now predict ahead
Reddit users spotted speculative decoding working on local GPUs—AI instantly outputting phrases via MTP. The local inference cost curve is being quiet
14h ago2 min read
Gemma 4llama.cpp
Fixing Gemma 4 Tool Calls in llama.cpp: Root Causes Explained
Four bugs in llama.cpp's Gemma 4 chat template handling caused tool call results to crash or loop.
Apr 82 min read