Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems
For network operators, this framework addresses the challenge of adapting physical layer systems to dynamic policies and real-time constraints, offering a significant performance gain.
The paper proposes Agentic-LTPO, a bilevel optimization framework using agentic AI to adapt physical layer configurations to evolving operator policies, achieving a 57.2% improvement in long-term performance over traditional methods in cell-free MIMO beamforming.
Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed objectives and constraints ineffective. This paper presents Agentic long-term performance optimization (Agentic-LTPO), a nested bilevel optimization framework that can be applied to adaptive physical layer problem configuration. The key idea is to employ agentic AI to generate upper-level configurations in a bilevel optimization structure, where evolving operator policies, environment summaries, and historical experiences are translated into structured lower-level optimization problem configurations. The lower level solves the problems with updated configurations for real-time physical-layer decisions. Considering cell-free MIMO beamforming as a use case, we embody Agentic-LTPO by designing a new multi-agent decision process with retrieval-augmented experience-based verification in the upper level, together with a closed-form beamformer in the lower level. Experiments demonstrate that Agentic-LTPO exhibits strong adaptability to dynamic operator policies and effectively enhances the system's long-term performance by 57.2% compared to traditional methods.