Back to home

Model Quantization

6 articles tagged with this topic

Quantization-Aware HealingLocalLLaMA

Quantization-Aware Healing: 4-bit Models Now Outperform Full-Precision Originals

r/LocalLLaMA study shows Quantization-Aware Healing lets 4-bit compressed LLMs outperform full-precision originals—potentially cutting deployment cost

4d ago2 min read
QwenRTX 3090

RTX 3090 Hits 82 tok/s on a 27B Model — Time to Retire the 'Cloud-Only' Myth

Dev squeezed a 27B Qwen model into 14GB VRAM on a 2020 RTX 3090 at 82 tokens/sec. Real signal: local AI hardware costs are now directly competing with

Aug 162 min read
LoRALLM Fine-Tuning

Updating 1% Params: Fine-Tuning & Quantization Slash Custom LLM Deployment Barriers

Fine-tuning turns LLMs into specialists; quantization trims them down. LoRA updates just 1% of params, enabling SMEs to customize AI with consumer GPU

May 52 min read
APEXQwen

APEX Quantizes 25 Models: 10B-Param AI on Home GPUs Flattens Compute Barrier

APEX quantizes 25+ MoE models with new I-Nano tier. 10B-param AI now runs on single consumer GPUs, slashing local deployment costs.

May 52 min read
QATModel Quantization

AI Quantization Ditches Full Downgrades for Mixed-Precision Topology

16-to-8-bit AI shifts crash precision. A new "equivalent topology" uses an 8-bit base, upgrading sensitive layers to 16-bit, balancing speed and preci

May 12 min read
QwenUnsloth

Qwen3.6-27B Quantized Fits Single Consumer GPU: Local Deployment Sweet Spot

Unsloth Q5-quantized Qwen3.6-27B runs stably on a single RTX 5090 across 19 rounds. Mid-size model local deployment is hitting the cost-capability swe

May 12 min read