AILOMAJun 29

HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial Observation

arXiv:2606.309666.5
Predicted impact top 78% in AI · last 90 daysOriginality Highly original
AI Analysis

This work addresses the challenge of specifying complex objectives and constraints in MARL, offering a more expressive alternative to reward shaping.

HyPOLE introduces a framework for multi-agent reinforcement learning under partial observability, guided by hyperproperty specifications in HyperLTL, and achieves clear advantages over baselines on SMAC, MessySMAC, and WildFire benchmarks.

Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathematical rigor; (2) expressiveness to specify objectives and constraints, and (3) the ability to define tactics to achieve objectives. However, these benefits remain largely unexplored in the context of Multi-Agent Reinforcement Learning (MARL). This paper introduces HyPOLE, a novel framework for MARL under partial observability, where learning is guided by the expressive power of the so-called hyperproperties and, in particular, the temporal logic HyperLTL. We integrate Centralized Training for Decentralized Execution (CTDE) techniques with HyPOLE to synthesize decentralized policies, and our evaluation on SMAC, MessySMAC, and WildFire benchmark demonstrates clear advantages over baselines.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes