What this is
Reddit user okoyl3 shared a set of benchmarks this week: after heavily modifying the open-source inference engine Strata (a program that efficiently packs model weights into hardware and generates text), he ran it on a 2018 IBM AC922 server — two POWER9 CPUs paired with four V100 GPUs. Combined with a quantized Qwen3.8-FN (compressing model parameters from high precision to low precision to save VRAM), prefill (when the model "reads" the input) peaked at 7,357 token/s, while decode (when the model "spits out" the output) hit roughly 113 token/s. The author says some of the changes will be contributed back to the community.
Industry view
Supporters argue this proves there's still significant headroom for inference optimization — that "old hardware plus good software" could substantially cut enterprise AI deployment costs, especially benefiting traditional industries sitting on POWER9 fleets. Skeptics counter: 113 token/s decode is only a small fraction of what an H100 delivers on similar models; NVLink 2.0's 70GB/s bandwidth will hit a wall fast; and this is a single-developer experiment, far from the stability, concurrency, and operability production environments demand. A demo of new uses for old hardware is not a shippable solution.
Impact on regular people
- For enterprise IT: Those sitting on old POWER9 clusters now have a "try before you replace GPUs" path — but don't rush to production. The gap between a single-dev demo and a production deployment spans stability, concurrency, and operations.
- For individual careers: No direct impact on daily workflows for now, but open-source inference engines are getting stronger; the hardware bar for running large models locally will keep falling.
- For consumer markets: No short-term impact; in the medium term, this could indirectly pressure down cloud AI inference pricing.