AIMay 14

A Survey on the Verification of Reinforcement Learning Policies

arXiv:2607.16210h-index: 39
Originality Synthesis-oriented
AI Analysis

For researchers and practitioners in safety-critical RL, this survey organizes a fragmented field, making it easier to understand and compare verification approaches.

This survey provides a unifying taxonomy and theoretical foundation for reinforcement learning policy verification, categorizing methods along three axes and identifying emerging directions.

Reinforcement learning (RL) is increasingly applied in complex, safety-critical domains, yet the lack of rigorous behavioral guarantees for neural network-based policies remains a major barrier to deployment. Recent advances in policy expressiveness and scale have intensified this challenge, leading to a rapidly growing but conceptually fragmented body of work on RL policy verification. This survey provides a unifying perspective on RL verification methods. We introduce a taxonomy that clarifies relationships among existing approaches along three axes: verification paradigm (formal versus probabilistic), temporal scope (step-wise versus multi-step), and guarantees strength. Beyond taxonomy, we unify underlying theoretical foundations, make implicit assumptions and limitations explicit, and identify emerging directions.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes