Back to home
Ninfer
2 articles tagged with this topic
NinferQwen3
Qwen3 Hits 220 tokens/sec on 5090 — Local AI Inflection Point Nears
New inference engine Ninfer pushed Qwen3 to 170 avg / 220 peak tokens/sec on RTX 5090 — over 2x faster than llama.cpp. Local AI is closing in on cloud
1d ago2 min read
NinferCMP170HX
AI as Programmer Completes CUDA Port — Mining GPU Doubles Qwen 35B Performance
AI coding agents helped a non-programmer port inference engine onto a mining GPU, doubling Qwen 35B. Mining cards + AI flatten the low-level tech bar.
Aug 222 min read