
NVIDIA DSX MaxLPS Lifts AI Factory Throughput 49% Within the Same Power Budget
A joint NVIDIA and Nscale test on GB300 NVL72 systems in Iceland ran 192 GPUs instead of 140 inside the same 264.4 kW budget, lifting normalized aggregate throughput 49.2% while per-instance speed barely moved.
NVIDIA has published a technical walkthrough of DSX MaxLPS, a power layer that shares power dynamically across GPUs in an AI data center instead of reserving a fixed peak for every node. The developer-blog article of September 27, 2026, by Sarah McKenney and Harry Petty, says operators can deploy up to 40% more GPUs within the same approved budget.
Why static provisioning strands power
An AI factory sits inside a hierarchy of electrical limits, from the utility connection down to racks, nodes and GPUs. Operators normally reserve enough for every accelerator to hit its peak at once, but real workloads do not behave that way: training alternates between compute, communication and checkpointing, while inference moves between prefill, decode, memory-bound work and idle periods. Under static per-node reservations, headroom inside one reservation cannot be handed to another node, and DSX MaxLPS answers this by monitoring actual consumption and reallocating power within operator-defined policies.

How the evaluation was run
Nscale deployed the software at its Verne campus data center in Keflavík, Iceland, while NVIDIA ran the workloads on GB300 NVL72 systems with Blackwell Ultra GPUs and Kimi K2.5 in FP4. The static baseline ran 140 GPUs; the MaxLPS setup ran 192, adding a third 52-GPU high-throughput instance under the same 264.4 kW budget.
What the measurements showed
Normalized aggregate throughput rose 49.2%, from 1,084,503 to 1,618,443 tokens per second, and throughput per provisioned watt climbed from 4.10 to 6.12 tokens/s/W. Per-instance output barely moved: 59,220 tokens/s for high-throughput instances against 59,153 in the baseline.
Mean GPU power went from 97.0 kW to 131.8 kW and power-budget utilization from 62.9% to 75.2%. Median and P75 latency stayed within 5% of baseline, but P99 time to first token worsened 17% from 15.7 seconds.
What operators should validate
NVIDIA recommends five stages: define the managed boundary and its enforceable limit, establish a representative baseline under static provisioning, introduce policies close to it, add capacity incrementally with testing at each step, and only then approve production limits.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.