Back to home
NInfer
3 articles tagged with this topic
NInfervLLM
Million-Token Context on Two 5090s — Amateur Dev Shatters Enterprise AI Myth
Reddit developer NInfer hits 1.04M token context on consumer RTX 5090s at 119 tok/s — 2.8x faster than vLLM with a 27B Qwen model.
2h ago2 min read
QwenRTX 4090
RTX 4090 Runs 27B Model Locally — On-Prem AI Hardware Cost Curve Breaks Through
A developer runs 27B Qwen on a single RTX 4090, hitting 250K-350K tokens context. Hardware bar for local LLMs slips below SME expectations.
Aug 162 min read
QwenNInfer
Qwen3.8 27B Runs 200 Tokens/Sec on a Single GPU, Closing Gap with Cloud APIs
Qwen3.8 27B with NInfer hits ~200 tokens/sec on a consumer RTX 5090. Local AI hardware barriers are falling fast, with real implications for enterpris
Aug 142 min read