Back to home
Inference Optimization
4 articles tagged with this topic
QwenRTX 3090
Two 3090s Run 27B Model at 165 tok/s — Local AI Is Finally 'Good Enough'
We noted a Reddit user hit 165 tok/s on a 27B Qwen model using two RTX 3090s (used rig under ¥20K) — Agent-ready. Local LLMs just crossed from 'toy' t
1d ago2 min read
Qwen3Nvidia
Qwen3 Hits 6250 token/s on RTX 5090: Open Source Drops Inference Costs Another 50%
Unsloth's compressed Qwen3 8B hits 6250 token/s on RTX 5090 — 50% faster than traditional Q4, powered by Nvidia's NVFP4 4-bit format.
Aug 212 min read
NVIDIANemotron
NVIDIA Compresses 66GB LLM to 22GB, 4x Faster — Inference Cost Story Rewritten
NVIDIA's Nemotron 3.5 Lightning gets NVFP4: 66GB to 22GB, 4x faster, near-lossless. The inference cost story just got rewritten — on-prem AI is now ch
Aug 172 min read
TurboQuantKV Cache
Independent KV Cache Evaluation SDK Signals Shift to Inference Infrastructure
KV cache dominates VRAM in long-context inference. An independent evaluation SDK for TurboQuant signals the shift from "can it run?" to "how to run st
May 52 min read