OCSYSYJun 22

Value iteration with stopping criterion: finite iterations, stability, and near-optimality guarantees

arXiv:2606.232293.1
Predicted impact top 72% in OC · last 90 daysOriginality Synthesis-oriented
AI Analysis

Provides a design framework for stopping criteria in VI that balances computational effort with stability and performance guarantees for control systems.

Value iteration (VI) is equipped with a generalized stopping criterion that ensures finite termination, stability, and explicit near-optimality bounds for deterministic discrete-time systems with infinite-horizon costs.

Value iteration (VI) is a cornerstone of dynamic programming that allows computing near-optimal feedback laws for general plant dynamics and cost functions. In practice, however, it must be stopped after finitely many iterations. This raises the question of when to stop the algorithm so that the resulting policies and value functions achieve desirable properties, like given near-optimality bounds and stability. In this context, we study deterministic, discrete-time systems with infinite-horizon (possibly discounted) costs whose inputs are generated by VI. We equip VI with a generalized stopping criterion that encompasses existing choices while allowing new ones. Our aim is to analyze the properties of the policies and value functions at the final iteration. Under mild assumptions, we first show that VI indeed terminates in a finite number of iterations. We then establish that the final policies are stabilizing by properly designing the stopping criterion, and derive explicit near-optimality bounds characterized by this choice. These results offer a design framework for the stopping criteria that balances computational effort with stability and performance guarantees.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes