This week, NVIDIA published DSX MaxLPS on its developer blog, redefining AI data centers' core metric from "how many GPUs do I have" to "how much AI can I run per kilowatt-hour." This signals NVIDIA publicly conceding — the AI industry's bottleneck is shifting from chips to power.
What this is
DSX MaxLPS is not a single product — it is a reference architecture for data center design. Its core question: when your grid capacity is fixed, how do you allocate that power to maximize AI inference output (the production phase where models actually serve end users)? The industry used to default to "more cards = stronger." Now we settle accounts by the watt. NVIDIA's answer: optimize power delivery, cooling, and scheduling together, and elevate "tokens per watt" (the basic unit of AI input/output) to the new KPI.
Industry view
The supportive case is straightforward: AI inference consumes a growing share of compute, and inference is extremely sensitive to per-watt performance. By promoting "compute per watt" to a KPI, NVIDIA is telling every AI data center operator that future competition is about power conversion efficiency — not rack count.
But there are dissenters. Framing the watt as the yardstick is, at its core, an NVIDIA commercial strategy: once GPUs are no longer the scarcest resource, the next round of competition becomes a system war over "power + cooling + software scheduling." This is NVIDIA expanding from "selling cards" to "selling the entire factory." Other analysts note that application-layer per-watt performance is highly workload-dependent with no universal optimum; an obsession with "per watt" could erode AI factory flexibility and cause operators to miss emerging workloads (e.g., AI Agents — systems where AI autonomously executes multi-step tasks).
Impact on regular people
For enterprise IT: When procuring AI inference services going forward, electricity costs may matter more than GPU procurement costs. When negotiating with cloud vendors, asking about "cost per watt per inference" is more practical than asking about "price per card-hour."
For individual careers: AI project evaluation is shifting from "is the model big enough" to "can the infrastructure hold up." Those who understand this will assess AI project ROI more accurately than their peers.
For the consumer market: If AI applications run on inefficient data centers, end users will eventually feel the impact through higher latency and rising prices — compute costs will pass through in some form.