We study diverse skill discovery in reward-free environments, aiming to discover all possible skills in simple grid-world environments where prior methods have struggled to succeed. This problem is formulated as mutual training of skills using an intrinsic reward and a discriminator trained to predict a skill given its trajectory. Our initial solution replaces the standard one-vs-all (softmax) discriminator with a one-vs-one (all pairs) discriminator and combines it with a novel intrinsic reward function and a dropout regularization technique. The combined approach is named APART: Diverse Skill Discovery using All Pairs with Ascending Reward and Dropout. We demonstrate that APART discovers all the possible skills in grid worlds with remarkably fewer samples than previous works. Motivated by the empirical success of APART, we further investigate an even simpler algorithm that achieves maximum skills by altering VIC, rescaling its intrinsic reward, and tuning the temperature of its softmax discriminator. We believe our findings shed light on the crucial factors underlying success of skill discovery algorithms in reinforcement learning.
翻译:摘要:本文研究无奖励环境中的多样化技能发现,旨在发现简单栅格世界环境中所有可能的技能——而此前方法在此类场景中难以成功。该问题被形式化为基于内在奖励的技能联合训练,并借助一个判别器根据轨迹预测对应技能。我们的初始方案将标准的一对多(softmax)判别器替换为一对一(全配对)判别器,并结合新颖的内在奖励函数与丢弃正则化技术。该方法被命名为APART:采用全配对、递增奖励与丢弃法实现多样化技能发现。实验表明,APART以远少于先前工作的样本数量,发现了栅格世界中的所有可能技能。受APART实证成功的启发,我们进一步研究了一种更简单的算法:通过调整VIC、重新缩放其内在奖励并调节其softmax判别器的温度参数,即可实现最大技能覆盖。我们相信,这些发现揭示了强化学习中技能发现算法成功的关键因素。