Recently, reinforcement learning has led to dexterous manipulation skills of increasing complexity. Nonetheless, learning these skills in simulation still exhibits poor sample-efficiency which stems from the fact these skills are learned from scratch without the benefit of any domain expertise. In this work, we aim to improve the sample efficiency of learning dexterous in-hand manipulation skills using controllers available via domain knowledge. To this end, we design simple sub-skill controllers and demonstrate improved sample efficiency using a framework that guides exploration toward relevant state space by following actions from these controllers. We are the first to demonstrate learning hard-to-explore finger-gaiting in-hand manipulation skills without the use of an exploratory reset distribution. Video results can be found at https://roamlab.github.io/vge
翻译:最近,强化学习推动了日益复杂灵巧操作技能的发展。然而,在仿真中学习这些技能仍存在样本效率低下的问题,其根源在于这些技能是从零开始学习,未借助任何领域知识。本研究旨在利用领域知识中的控制器来提高灵巧手内操作技能的学习样本效率。为此,我们设计了简单的子技能控制器,并通过一个框架遵循这些控制器的动作,将探索引导至相关状态空间,从而证明了样本效率的提升。我们是首个无需使用探索性重置分布,即可演示难以探索的手指步态手内操作技能学习的工作。视频结果详见 https://roamlab.github.io/vge