What this is

This week Alibaba's Qwen posted two numbers on NVIDIA's GB300 NVL72 rack (72 GPUs interconnected): 4,000 tokens per second per GPU, and 350 tokens per second at the user-facing endpoint. We read this as the clearest signal yet that inference marginal costs are being compressed faster than most forecasts suggested.

The Qwen3 2.4T parameter model shipped with same-day optimization for the new hardware — NVIDIA's official blog calls it "Day 0" support, no extra tuning required. To grasp what 350 tokens/sec feels like: the moment a user hits Enter, the model is already streaming words.

Industry view

The bull case: NVIDIA personally ran the benchmarks and co-published the post, which to us confirms deep silicon-and-model co-engineering between a US chipmaker and a Chinese frontier team. In our reading, the open-source community keeps pushing the ceiling on "domestic model + top-tier US compute" combos higher.

The bear case deserves recording. First, the 288k tokens/sec aggregate figure is FP8 precision (8-bit float weights, trading a sliver of accuracy for speed) — finance, medical, and other low-tolerance scenarios cannot drop this in directly. Second, a single GB300 rack carries a million-dollar-plus price tag; only hyperscalers and large clouds can afford to play, and the self-build bar for SMEs has been pushed into the stratosphere. Third, speed ≠ cheap — once power and depreciation are folded in, TCO is the real ledger. No matter how lively the open-source party gets, the check at the end is still signed by the same handful of hyperscale clouds.

Impact on regular people

  • For enterprise IT: the hardware bar for self-built inference clusters is rising from "hundreds of thousands" to "millions" within two years. Renting cloud capacity will increasingly beat buying machines for SMEs — we expect this trend to accelerate.
  • For working professionals: high-token tasks like long-document summarization and code generation will feel closer to "typing speed" in latency. Workflows that currently get interrupted mid-task will become smoother.
  • For consumer markets: pricing power for consumer AI subscriptions is getting squeezed. We expect 2026 to bring a price war across subscription products (ChatGPT Plus, Alibaba Tongyi, Wenxiaoyan, and others).