The ability to discover behaviours from past experience and transfer them to new tasks is a hallmark of intelligent agents acting sample-efficiently in the real world. Equipping embodied reinforcement learners with the same ability may be crucial for their successful deployment in robotics. While hierarchical and KL-regularized reinforcement learning individually hold promise here, arguably a hybrid approach could combine their respective benefits. Key to these fields is the use of information asymmetry across architectural modules to bias which skills are learnt. While asymmetry choice has a large influence on transferability, existing methods base their choice primarily on intuition in a domain-independent, potentially sub-optimal, manner. In this paper, we theoretically and empirically show the crucial expressivity-transferability trade-off of skills across sequential tasks, controlled by information asymmetry. Given this insight, we introduce Attentive Priors for Expressive and Transferable Skills (APES), a hierarchical KL-regularized method, heavily benefiting from both priors and hierarchy. Unlike existing approaches, APES automates the choice of asymmetry by learning it in a data-driven, domain-dependent, way based on our expressivity-transferability theorems. Experiments over complex transfer domains of varying levels of extrapolation and sparsity, such as robot block stacking, demonstrate the criticality of the correct asymmetric choice, with APES drastically outperforming previous methods.
翻译:从过往经验中发现行为并将其迁移至新任务的能力,是智能体在现实世界中实现样本高效决策的标志性特征。赋予具身强化学习智能体相同能力,对机器人领域的成功部署至关重要。尽管分层强化学习和KL正则化强化学习各自在此领域具有潜力,但混合方法有望结合两者的优势。这些方法的核心在于通过架构模块间的信息不对称性来偏置技能学习过程。虽然不对称性选择对迁移能力影响显著,但现有方法主要基于直觉进行领域无关的选择,这种选择方式可能并非最优。本文从理论和实验两方面证明了通过信息不对称性控制的技能在序列任务中关键的表达性-迁移性权衡。基于这一发现,我们提出面向表达性与可迁移技能的注意力先验(APES)——一种深度融合先验信息与层次结构的KL正则化分层方法。与现有方法不同,APES基于我们的表达性-迁移性定理,通过数据驱动的领域相关方式自动学习不对称性选择。在机器人叠方块等具有不同外推程度与稀疏性的复杂迁移域上的实验表明,正确选择不对称性至关重要,APES的性能显著超越先前方法。