Chunked prefill scheduling lowers the ramp rate of GPU power draw during LLM inference, not peak power, and this effect grows with server load and could reduce grid fast-ramping reserve needs by about 20%.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.SY 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Smoothing the Ramp, Not the Peak: Scheduling-Induced Power Dynamics of LLM Inference and Their Grid-Scale Consequences
Chunked prefill scheduling lowers the ramp rate of GPU power draw during LLM inference, not peak power, and this effect grows with server load and could reduce grid fast-ramping reserve needs by about 20%.