AI LG MLOct 5, 2020

Temporal Difference Uncertainties as a Signal for Exploration

Sebastian Flennerhag, Jane X. Wang, Pablo Sprechmann, Francesco Visin, Alexandre Galashov, Steven Kapturowski, Diana L. Borsa, Nicolas Heess, Andre Barreto, Razvan Pascanu

arXiv:2010.02255v215.816 citationsh-index: 60

Originality Highly original

AI Analysis

This addresses the problem of efficient exploration in non-tabular reinforcement learning for AI agents, offering a novel approach to uncertainty estimation that enhances exploration strategies.

The paper tackles the challenge of obtaining accurate uncertainty estimates for exploration in reinforcement learning with function approximators by proposing a method that induces a distribution over temporal difference errors to estimate value uncertainty, which is then used as an intrinsic reward to train a separate exploration policy, resulting in improved performance on hard exploration tasks like Deep Sea and Atari 2600 environments.

An effective approach to exploration in reinforcement learning is to rely on an agent's uncertainty over the optimal policy, which can yield near-optimal exploration strategies in tabular settings. However, in non-tabular settings that involve function approximators, obtaining accurate uncertainty estimates is almost as challenging a problem. In this paper, we highlight that value estimates are easily biased and temporally inconsistent. In light of this, we propose a novel method for estimating uncertainty over the value function that relies on inducing a distribution over temporal difference errors. This exploration signal controls for state-action transitions so as to isolate uncertainty in value that is due to uncertainty over the agent's parameters. Because our measure of uncertainty conditions on state-action transitions, we cannot act on this measure directly. Instead, we incorporate it as an intrinsic reward and treat exploration as a separate learning problem, induced by the agent's temporal difference uncertainties. We introduce a distinct exploration policy that learns to collect data with high estimated uncertainty, which gives rise to a curriculum that smoothly changes throughout learning and vanishes in the limit of perfect value estimates. We evaluate our method on hard exploration tasks, including Deep Sea and Atari 2600 environments and find that our proposed form of exploration facilitates both diverse and deep exploration.

View on arXiv PDF

Similar