Back to home
model compression
3 articles tagged with this topic
Qwenmodel compression
Qwen 4B Reasoning Jumps 17% With Almost No Size Increase
ByteOtter's QLAB method lifts Qwen 3.5 4B reasoning from 46.875 to 54.688 at IQ2_XS—a 16.67% gain with only 0.4% size increase.
Aug 222 min read
KLQquantization
One Person's Summer Project Cracks Quantization — Exposes the Decade-Long Blind Spot in Model
A solo researcher open-sourced KLQ, a training-free 4-bit quantization method that beats SpinQuant by measuring directional information density before
Aug 102 min read
NvidiaNVFP4
NVFP4 distillation hides internal geometry drift — speed gains mask structural damage
arXiv paper finds NVFP4 distillation preserves outputs but warps internal representations, hurting reasoning and coding.
Aug 92 min read