Pith. sign in

hub

Power stabilization for ai training datacenters

21 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

21 Pith papers citing it
1 external citations · Pith
abstract

Large Artificial Intelligence (AI) training workloads spanning several tens of thousands of GPUs present unique power management challenges. These arise due to the high variability in power consumption during the training. Given the synchronous nature of these jobs, during every iteration there is a computation-heavy phase, where each GPU works on the local data, and a communication-heavy phase where all the GPUs synchronize on the data. Because compute-heavy phases require much more power than communication phases, large power swings occur. The amplitude of these power swings is ever increasing with the increase in the size of training jobs. An even bigger challenge arises from the frequency spectrum of these power swings which, if harmonized with critical frequencies of utilities, can cause physical damage to the power grid infrastructure. Therefore, to continue scaling AI training workloads safely, we need to stabilize the power of such workloads. This paper introduces the challenge with production data and explores innovative solutions across the stack: software, GPU hardware, and datacenter infrastructure. We present the pros and cons of each of these approaches and finally present a multi-pronged approach to solving the challenge. The proposed solutions are rigorously tested using a combination of real hardware and Microsoft's in-house cloud power simulator, providing critical insights into the efficacy of these interventions under real-world conditions.

hub tools

citation-role summary

background 3

citation-polarity summary

years

2026 21

roles

background 3

polarities

support 2 background 1

representative citing papers

A Pre-Dispatch Resonance Safety Criterion for AI Training Clusters

eess.SY · 2026-06-20 · unverdicted · novelty 6.0

Derives a pre-dispatch resonance safety criterion by inverting two-area swing equations, bounding maximum safe AI cluster size at given iteration periods and showing rescheduling benefits on the IEEE 39-bus system.

Voltage Ride-Through in Large Loads- A Dual PQ Approach

eess.SY · 2026-05-01 · unverdicted · novelty 5.0

The paper proposes a dual PQ approach for voltage ride-through in large loads, showing that traditional reactive power compensation is limited by infrastructure constraints and that extreme dips may force disconnection.

Evaluating Grid Resilience in the Era of Ever-Increasing Data Centers

eess.SY · 2026-07-08 · conditional · novelty 3.0

Replacing a conventional load with energy-matched and scaled data center demand at a contingency-exposed bus in an IEEE 30-bus system increases unserved energy under transmission-constrained contingencies, with disruption-coincident demand amplifying the effect by 34.4%.

AI Infrastructure Sovereignty

cs.NI · 2026-02-11 · unverdicted · novelty 3.0

AI sovereignty requires coordinated design of data centers, optical networks, and real-time control systems to operate within energy availability and sustainability constraints.

citing papers explorer

Showing 21 of 21 citing papers.