LLM Inference
5 articles tagged with this topic
NVIDIA Slashes LLM Service Recovery to Seconds — But Only on Its GPUs
NVIDIA's Dynamo adds "Shadow Engine Recovery": crash recovery drops from minutes to seconds. Good news for enterprise AI — and a warning on GPU lock-i
Token Prices Dropped 90% in 2 Years — Here's Why LLM Inference Is Now Cheap
Why are some AI assistants free and others costly? We unpack Token, Transformer, and KV Cache—the hidden cost drivers behind every query.
Microsoft 4x LLM Inference: AI's Second Half Is Cutting Infra Costs
At NSDI 2026, Microsoft unveils AI infra breakthroughs like 4x LLM inference via cache sharing. AI competition shifts from scaling parameters to infra
80M Tokens for 4 RMB: DeepSeek Disk Cache Rewrites LLM Inference Costs
DeepSeek's novel architecture enables disk-level caching, slashing API costs 10x. This signals LLM inference shifting from raw compute to engineering
16 Nvidia DGX Spark Units Clustered for LLMs — Enterprise Compute Focus Shifts to VRAM
Reddit user clusters 16 Nvidia DGX Spark units, runs 434GB LLM. Unified memory validated. Inference bottlenecks shift from compute to VRAM — new path