A Predict-then-Schedule framework for Power Distribution Networks with AI Data Centers
This work addresses the problem of high energy costs and grid stress for power distribution networks serving AI data centers, offering an end-to-end learning approach that directly optimizes downstream scheduling performance.
The paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that integrates workload prediction with scheduling optimization for AI data centers, reducing operational cost and enhancing system security compared to conventional two-stage baselines.
The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing prediction error does not necessarily lead to minimized downstream operational loss. Hence, this paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that integrates upstream workload prediction with downstream scheduling optimization. By leveraging differentiable convex optimization, the PTS framework maps input features directly to optimal scheduling and enables gradient-based training. Furthermore, to respect the data center's capacity, a workload over-shifted loss combining electricity cost with a penalty for load-shedding is introduced to evaluate scheduling quality. Experiments demonstrate that the proposed framework significantly reduces operational cost and enhances system security compared to the conventional two-stage baseline.