GPU-to-Grid: Voltage Regulation via GPU Utilization Control
For data center operators and power grid engineers, this work provides a novel method to leverage LLM inference flexibility for voltage regulation, bridging a gap between computer and power systems research.
This paper introduces a GPU-to-Grid framework that uses GPU batch size control to regulate distribution-level voltage, demonstrating that adjusting GPU power can both alleviate and mitigate voltage violations, challenging the assumption that minimizing GPU power always benefits the grid.
While the rapid expansion of data centers poses challenges for power grids, it also offers new opportunities as flexible loads. Existing power system research often abstracts data centers as aggregate resources, while computer system research focuses on GPU energy efficiency and largely ignores grid impacts. To bridge this gap, we develop a GPU-to-Grid framework that couples device-level GPU control with power system objectives. We study distribution-level voltage regulation enabled by LLM inference flexibility, using batch size as a data-center-side control knob that trades off GPU power consumption, inference latency, and token throughput. We first formulate the problem as an optimization problem and then realize it as an online feedback optimization controller, implemented by the data center operator using its own empirical GPU power-performance model and real-time measurements from both the GPU and grid systems. Our key insight is that reducing GPU power alleviates lower-voltage violations, while increasing GPU power mitigates upper-voltage violations; this challenges the common belief that minimizing GPU power is always beneficial to power grids.