Back to home
inference engine
2 articles tagged with this topic
ZhipuTensorSharp
TensorSharp Doubles Llama.cpp Decoding Speed, Could Halve LLM Deployment Costs
GLM-5.3-Flash hits 2x llama.cpp decoding on new TensorSharp framework; local deployment costs may drop sharply, but ecosystem maturity is unproven.
1d ago2 min read
BitNetinference engine
BitNet hits 36 tokens/sec on a plain CPU — LLM inference starts shedding its GPU dependency
A developer shifu_legend wrote a zero-dependency inference engine in pure C99, hitting 36 tokens/sec on an Intel Xeon running a 1.58-bit BitNet model.
Aug 82 min read