DCAIAug 3

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling

arXiv:2608.018919.8
Predicted impact top 11% in DC · last 90 daysOriginality Highly original
AI Analysis

This work provides significant energy efficiency improvements for data center operators serving LLMs, addressing the growing energy consumption problem in AI infrastructure.

This paper tackles the problem of high energy consumption in large language model (LLM) serving by identifying that Attention and Feed-Forward Network (FFN) operators have different energy-optimal frequencies. The proposed AFlex framework, which jointly optimizes resource provisioning and GPU frequency scaling for disaggregated A/F serving, reduces energy per token by up to 49% compared to state-of-the-art disaggregated serving and 48% over frequency-scaling systems, while meeting service-level objectives.

Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption. Existing energy-management approaches adapt GPU frequencies only at the request or inference-phase level, overlooking operator-level differences in frequency sensitivity between Attention and feed-forward networks (FFNs). We find that the energy-optimal frequencies of Attention and FFN (A/F) differ and vary with the inference phase, workload, and system configurations. However, runtime variability and independent A/F frequency control create a large search space and high communication overhead. To address these challenges, we present AFlex, a framework that jointly optimizes resource provisioning and GPU frequency scaling for disaggregated A/F serving. AFlex introduces a global scheduler and a local operator-level dynamic voltage and frequency scaling (DVFS) controller to determine A/F resource allocations and frequencies. It further introduces an interleaved A/F pipeline with dynamic microbatch depth and adaptive request batching to reduce pipeline bubbles. We implement AFlex in SGLang and evaluate it on NVIDIA A800 GPUs using Qwen3-32B and Mixtral-8$\times$7B under production Conversation and Coding traces. \AFlex reduces energy per token by up to 49\% over state-of-the-art disaggregated serving and 48\% over frequency-scaling systems while satisfying TTFT and TPOT SLOs.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes