Back to home
Llama.cpp
4 articles tagged with this topic
Llama.cppMeta
Llama.cpp Hits 0.2 — The Bar for Running LLMs on Home PCs Drops Again
Open-source inference engine Llama.cpp releases 0.2.0, its first 0.2-series version, signaling a systemic architecture and performance overhaul worth
Aug 222 min read
QwenAlibaba Tongyi
Qwen 27B Runs 80-Step Agent on a Single GPU — Local Models Can Now Do Real Work
A Reddit test caught our eye: Qwen 27B on a consumer GPU made 80 autonomous tool calls from one prompt. The "cloud-only Agent" default is crumbling.
Aug 202 min read
AMDLlama.cpp
AMD GPUs Run Local LLMs 50% Faster — but Barriers Remain
Llama.cpp on ROCm 7.14 makes AMD Radeon 780M ~50% faster on dense models. AMD's first usable local AI option — limited to dense models, requires sourc
Aug 172 min read
Llama.cppGeorgi Gerganov
Llama.cpp's Gerganov Gets Collective Thanks: One Dev Holds Up Half of Open AI
Georgi Gerganov's Llama.cpp runs LLMs on ordinary laptops. Reddit's LocalLLaMA collectively thanked him—the entire local AI stack rests on his code.
Aug 162 min read