Humanoid locomotion is a key skill to bring humanoids out of the lab and into the real-world. Many motion generation methods for locomotion have been proposed including reinforcement learning (RL). RL locomotion policies offer great versatility and generalizability along with the ability to experience new knowledge to improve over time. This work presents a velocity-based RL locomotion policy for the REEM-C robot. The policy uses a periodic reward formulation and is implemented in Brax/MJX for fast training. Simulation results for the policy are demonstrated with future experimental results in progress.
翻译:人形机器人步态是实现其从实验室走向现实世界的关键能力。目前已提出包括强化学习在内的多种步态生成方法。强化学习步态策略具有出色的适应性和泛化能力,并能通过经验积累持续优化。本研究为REEM-C机器人提出了一种基于速度的强化学习步态策略。该策略采用周期性奖励机制,并基于Brax/MJX框架实现快速训练。文中展示了策略的仿真结果,实际机器人实验正在进行中。