Symmetry is a fundamental aspect of many real-world robotic tasks. However, current deep reinforcement learning (DRL) approaches can seldom harness and exploit symmetry effectively. Often, the learned behaviors fail to achieve the desired transformation invariances and suffer from motion artifacts. For instance, a quadruped may exhibit different gaits when commanded to move forward or backward, even though it is symmetrical about its torso. This issue becomes further pronounced in high-dimensional or complex environments, where DRL methods are prone to local optima and fail to explore regions of the state space equally. Past methods on encouraging symmetry for robotic tasks have studied this topic mainly in a single-task setting, where symmetry usually refers to symmetry in the motion, such as the gait patterns. In this paper, we revisit this topic for goal-conditioned tasks in robotics, where symmetry lies mainly in task execution and not necessarily in the learned motions themselves. In particular, we investigate two approaches to incorporate symmetry invariance into DRL -- data augmentation and mirror loss function. We provide a theoretical foundation for using augmented samples in an on-policy setting. Based on this, we show that the corresponding approach achieves faster convergence and improves the learned behaviors in various challenging robotic tasks, from climbing boxes with a quadruped to dexterous manipulation.
翻译:对称性是许多实际机器人任务中的一个基本特性。然而,当前的深度强化学习方法鲜少能够有效利用并发挥对称性的优势。学习到的行为往往未能实现期望的变换不变性,并受到运动伪影的困扰。例如,四足机器人在接受向前或向后运动指令时可能表现出不同的步态,尽管其身体关于躯干是对称的。这一问题在高维或复杂环境中变得更加突出,因为深度强化学习方法容易陷入局部最优,且无法均匀探索状态空间中的各个区域。以往在机器人任务中鼓励对称性的方法主要针对单任务场景进行探讨,此时对称性通常指运动本身的对称性,例如步态模式。本文重新审视这一主题,针对机器人领域中的目标条件任务,其中对称性主要存在于任务执行过程中,而非必然体现在学习到的运动本身。具体而言,我们研究了两种将对称不变性融入深度强化学习的方法——数据增强和镜像损失函数。我们为在策略设置中使用增强样本提供了理论基础。基于此,我们展示了相应方法如何在多种具有挑战性的机器人任务中(从四足机器人攀爬箱子到灵巧操作)实现更快的收敛速度并改进学习到的行为。