Back to home
model-quantization
2 articles tagged with this topic
tencentgguf
Tencent Compresses 1.5TB Model to 200GB — Local LLM Deployment Barrier Falls
Tencent compressed a 1.5TB LLM to 200GB, keeping ~98% performance. The on-prem hardware barrier is crossed — a subtle signal for cloud APIs.
9h ago2 min read
Gemmallama.cpp
Regular users now customize model files — local AI barrier drops a notch
Regular users now build GGUF files themselves on r/LocalLLaMA. Local AI is maturing — Chinese enterprises should rethink private deployment math.
Aug 222 min read