Back to home
CPU inference
2 articles tagged with this topic
BitNetinference engine
BitNet hits 36 tokens/sec on a plain CPU — LLM inference starts shedding its GPU dependency
A developer shifu_legend wrote a zero-dependency inference engine in pure C99, hitting 36 tokens/sec on an Intel Xeon running a 1.58-bit BitNet model.
Aug 82 min read
ggmlllama.cpp
GGML Adds Q1_0 1-Bit Quantization: Run 8B Models at 1.15GB
GGML now supports Q1_0 1-bit quantization, shrinking Bonsai 8B models to 1.15GB for CPU-only inference.
Apr 62 min read