Back to home
GGML
2 articles tagged with this topic
Llama.cppMeta
Llama.cpp Hits 0.2 — The Bar for Running LLMs on Home PCs Drops Again
Open-source inference engine Llama.cpp releases 0.2.0, its first 0.2-series version, signaling a systemic architecture and performance overhaul worth
Aug 222 min read
llama.cppQwen3
GPoUr with ~12gb vram and a 3080 getting 40tg/s on qwen3.6 35BA3B w/ 260k ctx
A llama.cpp fork with turbo3 KV cache quantization achieves ~40 tok/s on Qwen3-35 B-A3B with only 12GB VRAM.
Apr 162 min read