LG AIAug 24, 2023

APART: Diverse Skill Discovery using All Pairs with Ascending Reward and DropouT

Hadar Schreiber Galler, Tom Zahavy, Guillaume Desjardins, Alon Cohen

arXiv:2308.12649v12.0h-index: 25

Originality Incremental advance

AI Analysis

This addresses the challenge of efficiently discovering diverse skills in reinforcement learning for grid-world environments, representing an incremental improvement over existing methods.

The paper tackled the problem of diverse skill discovery in reward-free grid-world environments, where prior methods struggled, by introducing APART, which uses an all-pairs discriminator, novel intrinsic reward, and dropout to discover all possible skills with significantly fewer samples than previous works.

We study diverse skill discovery in reward-free environments, aiming to discover all possible skills in simple grid-world environments where prior methods have struggled to succeed. This problem is formulated as mutual training of skills using an intrinsic reward and a discriminator trained to predict a skill given its trajectory. Our initial solution replaces the standard one-vs-all (softmax) discriminator with a one-vs-one (all pairs) discriminator and combines it with a novel intrinsic reward function and a dropout regularization technique. The combined approach is named APART: Diverse Skill Discovery using All Pairs with Ascending Reward and Dropout. We demonstrate that APART discovers all the possible skills in grid worlds with remarkably fewer samples than previous works. Motivated by the empirical success of APART, we further investigate an even simpler algorithm that achieves maximum skills by altering VIC, rescaling its intrinsic reward, and tuning the temperature of its softmax discriminator. We believe our findings shed light on the crucial factors underlying success of skill discovery algorithms in reinforcement learning.

View on arXiv PDF

Similar