LGAIJul 1

Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL

arXiv:2607.003925.2
Predicted impact top 64% in LG · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in unsupervised RL, GenDa provides a more robust and data-efficient pre-training framework for skill-conditioned policies.

GenDa addresses non-stationary skill semantics and brittle generalization in unsupervised RL, achieving superior generalizability and data efficiency over prior methods.

Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks. Despite recent progress, we argue that current off-policy URL methods are limited by two critical, overlooked bottlenecks: (1) non-stationary skill semantics and (2) brittle generalization. To address these challenges, we propose GenDa (Generalizable Data-efficient Agent), a unified framework for robust unsupervised reinforcement learning. First, we introduce a skill relabeling mechanism to mitigate non-stationarity and significantly improve data efficiency for pre-training. Second, we propose a Complementary Information Bottleneck (CIB), encouraging the learned skill policy to focus on ego-centric features and become robust to distribution shifts for downstream tasks. Through various experiments, we demonstrate that GenDa significantly enhances the scalability of URL with superior generalizability and data efficiency. Our code and videos are available at https://ihatebroccoli.github.io/official-GenDa.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes