In this work, we propose algorithms and methods that enable learning dexterous object manipulation using simulated one- or two-armed robots equipped with multi-fingered hand end-effectors. Using a parallel GPU-accelerated physics simulator (Isaac Gym), we implement challenging tasks for these robots, including regrasping, grasp-and-throw, and object reorientation. To solve these problems we introduce a decentralized Population-Based Training (PBT) algorithm that allows us to massively amplify the exploration capabilities of deep reinforcement learning. We find that this method significantly outperforms regular end-to-end learning and is able to discover robust control policies in challenging tasks. Video demonstrations of learned behaviors and the code can be found at https://sites.google.com/view/dexpbt
翻译:本文提出了一套算法与方法,使配备多指手部末端执行器的单臂或双臂仿真机器人能够学习灵巧物体操作。借助并行GPU加速物理仿真器(Isaac Gym),我们为这些机器人实现了高难度任务,包括重新抓取、抓取投掷以及物体重定向。为解决这些问题,我们引入了一种去中心化的群体训练(PBT)算法,该算法可大幅增强深度强化学习的探索能力。实验表明,该方法显著优于常规端到端学习,能够在高难度任务中发现鲁棒的控制策略。学习行为演示视频及代码见https://sites.google.com/view/dexpbt