AI Alignment through a Game-theoretic Lens: A Survey
This survey provides a structured overview for AI alignment researchers, highlighting where game theory can genuinely improve the robustness and adaptability of AI systems.
This survey examines AI alignment through a game-theoretic lens, organizing recent progress around key game-theoretic elements. It synthesizes the literature along three challenges: preference diversity, alignment priority, and temporal dynamics, clarifying the benefits and limitations of game theory in current alignment methods.
As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessness, and controllability, often struggle to capture real-world preferences that are context-dependent, non-transitive, and shaped by dynamic multi-party interactions. This survey reviews AI alignment through a game-theoretic lens. Specifically, it organizes recent progress around key game-theoretic elements and synthesizes the literature along three challenges: preference diversity, alignment priority, and temporal dynamics. This perspective clarifies where current alignment methods genuinely benefit from game-theoretic analysis, where the framework is looser, and what challenges remain in building robust, adaptive, and verifiable AI systems.