LGMay 22, 2025

Reward-Aware Proto-Representations in Reinforcement Learning

Hon Tik Tse, Siddarth Chandrasekar, Marlos C. Machado

arXiv:2505.16217v113.05 citationsh-index: 22

Originality Incremental advance

AI Analysis

This work addresses a key challenge in reinforcement learning for researchers and practitioners by extending representation learning to be reward-aware, though it is incremental as it builds on the SR framework.

The paper tackles the limitation of the successor representation (SR) in reinforcement learning by proposing the default representation (DR), which incorporates reward dynamics, and shows that the DR leads to qualitatively different behavior and quantitatively better performance in tasks like reward shaping and exploration.

In recent years, the successor representation (SR) has attracted increasing attention in reinforcement learning (RL), and it has been used to address some of its key challenges, such as exploration, credit assignment, and generalization. The SR can be seen as representing the underlying credit assignment structure of the environment by implicitly encoding its induced transition dynamics. However, the SR is reward-agnostic. In this paper, we discuss a similar representation that also takes into account the reward dynamics of the problem. We study the default representation (DR), a recently proposed representation with limited theoretical (and empirical) analysis. Here, we lay some of the theoretical foundation underlying the DR in the tabular case by (1) deriving dynamic programming and (2) temporal-difference methods to learn the DR, (3) characterizing the basis for the vector space of the DR, and (4) formally extending the DR to the function approximation case through default features. Empirically, we analyze the benefits of the DR in many of the settings in which the SR has been applied, including (1) reward shaping, (2) option discovery, (3) exploration, and (4) transfer learning. Our results show that, compared to the SR, the DR gives rise to qualitatively different, reward-aware behaviour and quantitatively better performance in several settings.

View on arXiv PDF

Similar