Model-free reinforcement learning is a promising approach for autonomously solving challenging robotics control problems, but faces exploration difficulty without information of the robot's kinematics and dynamics morphology. The under-exploration of multiple modalities with symmetric states leads to behaviors that are often unnatural and sub-optimal. This issue becomes particularly pronounced in the context of robotic systems with morphological symmetries, such as legged robots for which the resulting asymmetric and aperiodic behaviors compromise performance, robustness, and transferability to real hardware. To mitigate this challenge, we can leverage symmetry to guide and improve the exploration in policy learning via equivariance/invariance constraints. In this paper, we investigate the efficacy of two approaches to incorporate symmetry: modifying the network architectures to be strictly equivariant/invariant, and leveraging data augmentation to approximate equivariant/invariant actor-critics. We implement the methods on challenging loco-manipulation and bipedal locomotion tasks and compare with an unconstrained baseline. We find that the strictly equivariant policy consistently outperforms other methods in sample efficiency and task performance in simulation. In addition, symmetry-incorporated approaches exhibit better gait quality, higher robustness and can be deployed zero-shot in real-world experiments.
翻译:无模型强化学习是一种有望自主解决具有挑战性的机器人控制问题的方法,但在缺乏机器人运动学和动力学形态信息时面临探索困难。对称状态下的多模态探索不足导致行为往往不自然且次优。这一问题在具有形态对称性的机器人系统中尤为突出,例如腿式机器人,其由此产生的不对称和非周期性行为会损害性能、鲁棒性以及向真实硬件的迁移能力。为缓解这一挑战,我们可以利用对称性通过等变/不变约束来引导和改进策略学习中的探索。本文研究了两种融入对称性方法的有效性:修改网络架构使其严格等变/不变,以及利用数据增强来近似等变/不变的行动者-评论家模型。我们在具有挑战性的操控与双足运动任务上实现这些方法,并与无约束基线进行比较。结果表明,严格等变策略在仿真中的样本效率和任务性能上始终优于其他方法。此外,融入对称性的方法展现出更好的步态质量、更高的鲁棒性,并且可以在真实实验中实现零样本部署。