LGDec 22, 2016

On the function approximation error for risk-sensitive reinforcement learning

arXiv:1612.07562v151.0

Originality Synthesis-oriented

AI Analysis

This work provides incremental improvements in error analysis for policy evaluation in risk-sensitive RL, benefiting researchers in reinforcement learning theory.

The authors tackled the problem of bounding function approximation error in risk-sensitive reinforcement learning using exponential utility, obtaining new error bounds that achieve the actual error in examples where prior bounds were weaker.

In this paper we obtain several informative error bounds on function approximation for the policy evaluation algorithm proposed by Basu et al. when the aim is to find the risk-sensitive cost represented using exponential utility. The main idea is to use classical Bapat's inequality and to use Perron-Frobenius eigenvectors (exists if we assume irreducible Markov chain) to get the new bounds. The novelty of our approach is that we use the irreduciblity of Markov chain to get the new bounds whereas the earlier work by Basu et al. used spectral variation bound which is true for any matrix. We also give examples where all our bounds achieve the "actual error" whereas the earlier bound given by Basu et al. is much weaker in comparison. We show that this happens due to the absence of difference term in the earlier bound which is always present in all our bounds when the state space is large. Additionally, we discuss how all our bounds compare with each other. As a corollary of our main result we provide a bound between largest eigenvalues of two irreducibile matrices in terms of the matrix entries.

View on arXiv PDF

Similar