The past decade has seen vast progress in deep reinforcement learning (RL) on the back of algorithms manually designed by human researchers. Recently, it has been shown that it is possible to meta-learn update rules, with the hope of discovering algorithms that can perform well on a wide range of RL tasks. Despite impressive initial results from algorithms such as Learned Policy Gradient (LPG), there remains a generalization gap when these algorithms are applied to unseen environments. In this work, we examine how characteristics of the meta-training distribution impact the generalization performance of these algorithms. Motivated by this analysis and building on ideas from Unsupervised Environment Design (UED), we propose a novel approach for automatically generating curricula to maximize the regret of a meta-learned optimizer, in addition to a novel approximation of regret, which we name algorithmic regret (AR). The result is our method, General RL Optimizers Obtained Via Environment Design (GROOVE). In a series of experiments, we show that GROOVE achieves superior generalization to LPG, and evaluate AR against baseline metrics from UED, identifying it as a critical component of environment design in this setting. We believe this approach is a step towards the discovery of truly general RL algorithms, capable of solving a wide range of real-world environments.
翻译:过去十年间,在人类研究人员手工设计的算法推动下,深度强化学习取得了巨大进展。近期研究表明,可以通过元学习来更新规则,以期发现能在广泛强化学习任务中表现优异的算法。尽管学习策略梯度等算法已取得令人瞩目的初步成果,但当这些算法应用于未见环境时仍存在泛化差距。本研究探讨了元训练分布特征如何影响此类算法的泛化性能。受此分析启发,并基于无监督环境设计的思想,我们提出一种自动生成课程的新方法,旨在最大化元学习优化器的遗憾值,同时提出一种新的遗憾近似方法,称之为算法遗憾。由此产生的方法被命名为通过环境设计获得的通用强化学习优化器(GROOVE)。系列实验表明,GROOVE相比学习策略梯度具有更优的泛化能力,我们还将算法遗憾与无监督环境设计的基线指标进行对比,证实其在该场景环境设计中的关键作用。我们认为该方法向发现能够解决广泛真实环境问题的真正通用强化学习算法迈出了重要一步。