Cloud Services

Lambda Maximizes Performance per Watt With NVIDIA DSX MaxLPS

Objective

As AI workloads scale and compute demand outpaces available infrastructure, fixed power budgets have emerged as a critical bottleneck, constraining both throughput and performance per watt

To maximize energy efficiency, Lambda is leveraging the NVIDIA DSX™ AI Factory Platform (NVIDIA DSX) which unifies design, simulation, and operations for maximum efficiency. As one of the first cloud providers to use NVIDIA DSX MaxLPS™—a core component of NVIDIA DSX—on NVIDIA HGX B200 GPU servers, Lambda is unlocking greater cluster performance and efficiency.

Customer

Lambda

Use Case

Data Center / Cloud

Key Takeaways

+24% Cluster-Wide Token Throughput

  • Using NVIDIA DSX MaxLPS, Lambda ran 19 nodes instead of 16 within the same power budget, increasing cluster-wide token throughput by 24% and GPU capacity by 19%.

+23% Performance per Watt

  • NVIDIA DSX MaxLPS increased cluster-wide throughput per watt by 23% while operating within the same fixed power budget.

+20% Hybrid Inference Throughput

  • NVIDIA DSX MaxLPS enabled Lambda to add two inference nodes within the same power budget, increasing throughput by 20% in concurrent inference and training configurations.

+17% Hybrid Training Throughput

  • NVIDIA DSX MaxLPS enabled Lambda to add two training nodes within the same power budget, increasing hybrid training-cluster throughput by 17%.

Power Is the AI Factory’s Limiting Constraint

Lambda was founded in 2012 by published machine learning engineers, who started by building AI applications and realized that scaling was bottlenecked by infrastructure. In 2018, Lambda became one of the first AI cloud providers to deploy NVIDIA accelerated computing in the cloud. Today, Lambda serves over ten thousand customers, from academia to AI labs to enterprise AI teams to hyperscalers, and its fleet of NVIDIA AI infrastructure includes the NVIDIA Ampere, NVIDIA Hopper™, NVIDIA Blackwell, and NVIDIA Vera Rubin platforms.

Deploying and managing large-scale clusters for a wide variety of customers is Lambda’s core business and its most demanding operating challenge. AI factories are purpose-built to manufacture intelligence, where every kilowatt counts toward delivering tokens per second. At production scale, facility power budgets are fixed. NVIDIA accelerated computing systems are engineered to run the most intensive AI workloads at maximum throughput per watt. When demand exceeds the compute capacity, the question shifts from hardware procurement to operational control—how many nodes can run productively within the fixed power budget. 

Cost per token is defined by a simple fraction: infrastructure cost per GPU-hour divided by tokens delivered. Stranded power and static configurations reduce that denominator. Without real-time power shaping, cloud operators face a blunt choice: provision conservatively for peak draw and leave capacity idle, or pack in more nodes and risk instability when workloads spike.

Specialized clouds like Lambda need a mechanism to monitor and allocate power dynamically at the rack level, bringing more nodes online within an existing power budget without sacrificing reliability or throughput.

DSX MaxLPS is designed to maximize compute within a fixed power budget. Lambda’s results demonstrate its immediate impact: MaxLPS enabled 19 nodes to run within the power budget of a 16-node baseline, increasing cluster token throughput by approximately 24% and performance per watt by 23%.

Lambda validated DSX MaxLPS across a five-rack,19-node cluster of HGX B200 systems and using MLPerf inference and training workloads to generate consistent, peak-level power draw and reproducible, industry-comparable results.

  • Dynamic Power Software (DPS) monitors GPU and rack-level consumption in real time, reallocating power to where it is needed to maximize token throughput within fixed budgets.
  • Configurable per-policy power controls let operators set power targets per node without changes to running workloads; Lambda tested 80% and 85% of maximum draw, reducing per-node consumption enough to bring additional nodes online within the same facility budget.
  • Power-observability dashboards give operators visibility into rack-level draw—a capability refined through Lambda’s production feedback.

Inference Throughput (Tokens/s): GPT-OSS-120B, 40 QPS per node 

More Performance, Same Power Budget

The results demonstrated that DSX MaxLPS meaningfully expands token throughput within a fixed power envelope and that the gains compound when mixed workloads create natural headroom for dynamic power sharing. By interleaving the bursty power profiles of training workloads with the steadier demand of inference, this approach recovers capacity that would otherwise remain stranded in statically provisioned clusters.

For pure inference, the results are direct:

  • Baseline (16 nodes, no policy): ~4.04M tokens/sec
  • 85% DSX MaxLPS policy (19 nodes): ~5M tokens/sec—a 24% gain in cluster token performance 

The concurrent workload results reveal an equally important dynamic. Production AI factories rarely run a single workload type—training and inference run side by side, each with different power profiles. Under DSX MaxLPS, that heterogeneity becomes an efficiency asset. Training draws power in bursts; inference fills the gaps. DSX MaxLPS allocates power dynamically across the mix, unlocking stranded capacity that static configurations leave idle. At the 80% policy running 10 inference nodes and 10 training nodes simultaneously, training cluster throughput increased 17% and inference cluster throughput rose 20%. Mixed workloads highlight the capacity optimization potential of DSX MaxLPS, though production implementation requires careful tuning to balance throughput gains with latency and stability requirements.

Beyond throughput, Lambda’s collaboration with NVIDIA’s engineering teams contributed to improved usability and configurable power-shaping controls that will be available within DSX MaxLPS. 

Lambda’s adoption establishes a blueprint for the next generation of AI factories: Maximizing performance per watt—not raw hardware provisioning—is the key lever for improving AI factory economics and delivering lowest token costs.

“With our proof of concept, we believe we’ve moved beyond the limitation of fixed power budgets. NVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint. It’s a scalable and sustainable blueprint for how we can operate and expand our infrastructure moving forward.”

Dave Ward
President, Cloud Services, Lambda

Extending DSX MaxLPS to Next-Generation Infrastructure

As the collaboration continues, Lambda may be well positioned to deploy DSX MaxLPS across its diverse and expansive fleet as plans and timelines become clearer—giving customers more token capacity per rack, more precise control over power allocation, and a model for sustainable AI factory operations that the rest of the industry can follow.

Explore how NVIDIA DSX MaxLPS is enabling greater performance per watt.

Related Customer Stories