NVFP4
6 articles tagged with this topic
NVIDIA's NVFP4 Doubles 5090 Speed — But Quality Debate Won't Die
NVIDIA's NVFP4 4-bit quantization on RTX 5090 promises 2x local AI inference speed, but a 200+ reply Reddit thread shows users can't agree on quality
Local 8B Model Ties 27B on Agent Coding — The 'Good Enough' Moment Arrives
Reddit dev's local Agent coding benchmark on RTX Pro 6000: 8B quantized model nearly matches 27B. Hardware math may need rewriting.
16GB GPU Hits 110 token/s — Local AI Finally Matches Cloud Speed
FlashML's open-source FreeToken tool hits 110 token/s on a Qwen3 35B model using a 16GB consumer GPU. If reproducible, local AI could start rivaling c
NVIDIA Compresses 66GB LLM to 22GB, 4x Faster — Inference Cost Story Rewritten
NVIDIA's Nemotron 3.5 Lightning gets NVFP4: 66GB to 22GB, 4x faster, near-lossless. The inference cost story just got rewritten — on-prem AI is now ch
NVFP4 distillation hides internal geometry drift — speed gains mask structural damage
arXiv paper finds NVFP4 distillation preserves outputs but warps internal representations, hurting reasoning and coding.
NVIDIA NVFP4 Puts 26B Model on Consumer GPU With Under 1% Accuracy Loss
NVIDIA's NVFP4 Gemma-4-26B shrinks to 18.8GB for consumer GPUs with <0.7% accuracy loss. 4-bit is now optimal, but also an ecosystem lock-in.