LGJun 18

Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning

arXiv:2606.204117.9
Predicted impact top 56% in LG · last 90 daysOriginality Incremental advance
AI Analysis

Improves sample efficiency and scalability of DAE-based RL for practitioners dealing with high-dimensional observations and partial observability.

Extended Direct Advantage Estimation to partially observable domains and reduced its computational overhead using discrete latent dynamics, achieving scalable and sample-efficient deep reinforcement learning on Atari games.

Direct Advantage Estimation (DAE) has been shown to improve the sample efficiency of deep reinforcement learning algorithms. However, its reliance on full environment observability limits its applicability in realistic settings, and its requirement to model transition probabilities incurs substantial computational overhead for high-dimensional observations. In the present work, we address both limitations. First, we extend the theoretical framework of DAE to partially observable domains with minimal modifications. Second, we reduce its computational complexity by introducing discrete latent dynamics models that efficiently approximate transition probabilities. We evaluate our approach on the Arcade Learning Environment and find that DAE scales effectively with function approximator capacity while retaining high sample efficiency.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes