Smoothing the Ramp, Not the Peak: Scheduling-Induced Power Dynamics of LLM Inference and Their Grid-Scale Consequences
Pan Li, Yize Chen, Xia Miao, Dai Wang
Abstract
Large language model (LLM) inference serving is a fast-growing electricity load whose power dynamics remain uncharacterized from a grid-planning perspective. Using real, measured GPU power traces, we show that chunked prefill scheduling, a latency-motivated technique already deployed by default in production LLM serving, is a controllable knob that regulates power ramp rate without touching peak power. Contrary to the intuitive hypothesis that splitting a long prompt's computation into smaller steps should flatten its power spike, peak power stays relatively the same while mean ramp rate falls substantially. Critically, this ramp-rate benefit is not a fixed property of the policy: it grows monotonically with system saturation, and we confirm this along two independent axes: concurrency (7.0% at light load to 34.6% at heavy load, mean-ramp reduction) and long-prompt ("whale") request load (from statistically flat at low whale incidence to 42.6% at high whale fraction/size). We translate this single-GPU mechanism into an operational grid quantity, regulation-reserve procurement, posed and solved as a chance-constrained problem using a model-free bootstrap directly resampling real measured power traces. At a representative operating point, this translates to an estimated 20.3-22.7% reduction in the fast-ramping reserve capacity a grid operator would need to provision, across reliability levels from 95% to 99.9%. Together, these results give grid operators a no-cost demand-shaping tool available today, whose benefit is largest precisely when data centers run hottest and grid stress is most salient.
Create a lesson
Related papers
Leader-Follower Formation Control with Prescribed Convergence Rates under Bearing Persistence of Excitation
Tarek Bouazza, Zhiqi Tang, Soulaimane Berkane et al.
On asymptotic stability of the time-varying Kalman filter for unstabilizable linear systems: an optimization perspective
James B. Rawlings, Titus Quah, Matthias A. Müller
Designing Grid-Aware Dynamic Specifications for Large Data Center Loads
Ashutossh Gupta, Vassilis Kekatos
Time-Optimal Operation of a Load-Hoisting Gantry Crane
Eric Mountain, Tarunraj Singh
Learning to Solve Two-Stage Stochastic Unit Commitment Problems with Quality Guarantees
Andrea Fusco, Andrea Lodi, Lavanya Marla
Towards Interaction Regulation from Human Feedback via Free Energy Minimization
Maria Paula Diaz Monfort, Cinzia Tomaselli, Michael Richardson et al.