Cloud Services
As AI workloads scale and compute demand outpaces available infrastructure, fixed power budgets have emerged as a critical bottleneck, constraining both throughput and performance per watt.
To maximize energy efficiency, Lambda is leveraging the NVIDIA DSX™ AI Factory Platform (NVIDIA DSX) which unifies design, simulation, and operations for maximum efficiency. As one of the first cloud providers to use NVIDIA DSX MaxLPS™—a core component of NVIDIA DSX—on NVIDIA HGX B200 GPU servers, Lambda is unlocking greater cluster performance and efficiency.
+24% Cluster-Wide Token Throughput
+23% Performance per Watt
+20% Hybrid Inference Throughput
+17% Hybrid Training Throughput
Lambda was founded in 2012 by published machine learning engineers, who started by building AI applications and realized that scaling was bottlenecked by infrastructure. In 2018, Lambda became one of the first AI cloud providers to deploy NVIDIA accelerated computing in the cloud. Today, Lambda serves over ten thousand customers, from academia to AI labs to enterprise AI teams to hyperscalers, and its fleet of NVIDIA AI infrastructure includes the NVIDIA Ampere, NVIDIA Hopper™, NVIDIA Blackwell, and NVIDIA Vera Rubin platforms.
Deploying and managing large-scale clusters for a wide variety of customers is Lambda’s core business and its most demanding operating challenge. AI factories are purpose-built to manufacture intelligence, where every kilowatt counts toward delivering tokens per second. At production scale, facility power budgets are fixed. NVIDIA accelerated computing systems are engineered to run the most intensive AI workloads at maximum throughput per watt. When demand exceeds the compute capacity, the question shifts from hardware procurement to operational control—how many nodes can run productively within the fixed power budget.
Cost per token is defined by a simple fraction: infrastructure cost per GPU-hour divided by tokens delivered. Stranded power and static configurations reduce that denominator. Without real-time power shaping, cloud operators face a blunt choice: provision conservatively for peak draw and leave capacity idle, or pack in more nodes and risk instability when workloads spike.
Specialized clouds like Lambda need a mechanism to monitor and allocate power dynamically at the rack level, bringing more nodes online within an existing power budget without sacrificing reliability or throughput.
DSX MaxLPS is designed to maximize compute within a fixed power budget. Lambda’s results demonstrate its immediate impact: MaxLPS enabled 19 nodes to run within the power budget of a 16-node baseline, increasing cluster token throughput by approximately 24% and performance per watt by 23%.
Lambda validated DSX MaxLPS across a five-rack,19-node cluster of HGX B200 systems and using MLPerf inference and training workloads to generate consistent, peak-level power draw and reproducible, industry-comparable results.
Inference Throughput (Tokens/s): GPT-OSS-120B, 40 QPS per node
The results demonstrated that DSX MaxLPS meaningfully expands token throughput within a fixed power envelope and that the gains compound when mixed workloads create natural headroom for dynamic power sharing. By interleaving the bursty power profiles of training workloads with the steadier demand of inference, this approach recovers capacity that would otherwise remain stranded in statically provisioned clusters.
For pure inference, the results are direct:
The concurrent workload results reveal an equally important dynamic. Production AI factories rarely run a single workload type—training and inference run side by side, each with different power profiles. Under DSX MaxLPS, that heterogeneity becomes an efficiency asset. Training draws power in bursts; inference fills the gaps. DSX MaxLPS allocates power dynamically across the mix, unlocking stranded capacity that static configurations leave idle. At the 80% policy running 10 inference nodes and 10 training nodes simultaneously, training cluster throughput increased 17% and inference cluster throughput rose 20%. Mixed workloads highlight the capacity optimization potential of DSX MaxLPS, though production implementation requires careful tuning to balance throughput gains with latency and stability requirements.
Beyond throughput, Lambda’s collaboration with NVIDIA’s engineering teams contributed to improved usability and configurable power-shaping controls that will be available within DSX MaxLPS.
Lambda’s adoption establishes a blueprint for the next generation of AI factories: Maximizing performance per watt—not raw hardware provisioning—is the key lever for improving AI factory economics and delivering lowest token costs.
“With our proof of concept, we believe we’ve moved beyond the limitation of fixed power budgets. NVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint. It’s a scalable and sustainable blueprint for how we can operate and expand our infrastructure moving forward.”
Dave Ward
President, Cloud Services, Lambda
As the collaboration continues, Lambda may be well positioned to deploy DSX MaxLPS across its diverse and expansive fleet as plans and timelines become clearer—giving customers more token capacity per rack, more precise control over power allocation, and a model for sustainable AI factory operations that the rest of the industry can follow.
Explore how NVIDIA DSX MaxLPS is enabling greater performance per watt.