What this is

This week, a post on r/LocalLLaMA put hard numbers on the table: running a 70GB model on a 64GB machine with Strata, a new inference engine, pushed generation speed from 21 t/s (tokens per second) to 60 t/s—but memory usage hit 96%, and the user couldn't even launch image generation.

Poster Cautious_Chicken_604 runs an AMD R9700 + RTX 5060 Ti + 64GB DDR5 setup. He normally uses the R9700 to run Qwen3-27B (Q6 quantization—compressing model weights to 6-bit precision) as an assistant, and the 5060 Ti for AI image generation in ComfyUI. He wanted to upgrade to a larger Qwen3.8-Flash-Next, which under llama.cpp (the dominant open-source inference engine) crawled at 21 t/s, but jumped to 60 t/s after switching to Strata. With the model loaded, memory sat at 96%; opening ComfyUI froze the system dead.

His own summary: feels blessed (finally runs epic-tier models), and cursed (LLM or image generation—pick one).

Industry view

We see two things here. First, the kind of "inference engine optimization" Strata represents is advancing fast: same machine, same model, speed jumped from 21 t/s to 60 t/s. A pure software win. Second, the hardware ceiling hasn't moved: a 70GB model eats 70GB of memory, and no clever engine can route around that.

The counterargument is equally clear: this is fundamentally an enthusiast's complaint. Ordinary users pay $20/month for ChatGPT or Claude and never bump into memory ceilings. Cloud APIs (calling AI models over the network) are the answer for 99% of people; local deployment is closer to a specialty battlefield for privacy-sensitive industries—finance, healthcare, legal. And the Qwen team keeps shipping larger models, so hardware anxiety will only intensify.

One more detail: the user admits that after this year's hardware splurge, his wife demanded "hedging" via jewelry and international travel—so he can't upgrade to 128GB RAM for now. The local AI hardware arms race is already spilling into household budgets.

Impact on regular people

For enterprise IT: Total cost of ownership for on-prem AI is being underestimated. Memory expansion kits, like high-end GPUs, are a major CapEx (capital expenditure) line item. When evaluating local AI, enterprises must budget for memory upgrades over the next three years.

For individual professionals: Running local AI on a company machine to safeguard client data requires a higher hardware bar than most expect. 64GB is just entry-level; 128GB is the comfort zone—and bulk procurement isn't cheap.

For the consumer market: DIY AI workstations are splitting into "VRAM camps" and "RAM camps," complicating hardware buying. Casual users no longer face just "which GPU?" but the whole stack: GPU, memory, and inference engine, matched together.