Learning-based grasping can afford real-time grasp motion planning of multi-fingered robotics hands thanks to its high computational efficiency. However, learning-based methods are required to explore large search spaces during the learning process. The search space causes low learning efficiency, which has been the main barrier to its practical adoption. In addition, the trained policy lacks a generalizable outcome unless objects are identical to the trained objects. In this work, we develop a novel Physics-Guided Deep Reinforcement Learning with a Hierarchical Reward Mechanism to improve learning efficiency and generalizability for learning-based autonomous grasping. Unlike conventional observation-based grasp learning, physics-informed metrics are utilized to convey correlations between features associated with hand structures and objects to improve learning efficiency and outcomes. Further, the hierarchical reward mechanism enables the robot to learn prioritized components of the grasping tasks. Our method is validated in robotic grasping tasks with a 3-finger MICO robot arm. The results show that our method outperformed the standard Deep Reinforcement Learning methods in various robotic grasping tasks.
翻译:基于学习的抓取方法因其高计算效率,能够实现多指机器人手的实时抓取运动规划。然而,基于学习的方法在学习过程中需要探索较大的搜索空间。该搜索空间导致学习效率低下,这已成为其实际应用的主要障碍。此外,除非物体与训练物体完全相同,否则训练得到的策略缺乏泛化能力。本文提出了一种新颖的物理引导深度强化学习与分层奖励机制,以提高基于学习的自主抓取的学习效率和泛化能力。与传统的基于观测的抓取学习不同,本文利用物理信息度量来传递与手部结构和物体相关的特征之间的相关性,以提高学习效率和结果。此外,分层奖励机制使机器人能够学习抓取任务中的优先组件。我们的方法在三指MICO机械臂的机器人抓取任务中得到了验证。结果表明,在各种机器人抓取任务中,我们的方法优于标准的深度强化学习方法。