Achieving expert-level expressive full-body motion tracking across multiple humanoids solely from demonstration data remains a challenging and relatively an underexplored problem in humanoid robot learning. Cross-embodiment motion tracking policies are mostly trained by decoupling the control problem into upper and lower body control. This work proposes VENOM, a cross-embodiment full-body motion tracking model for humanoids in simulation. VENOM is a GPT-based motion tracker trained on multiple humanoid data that can track the entire body without the requirement to split into upper and lower body control. We curate a multi-humanoid motion tracking dataset called the VENOM dataset that contains states, actions, and rewards and train VENOM and the baselines on this dataset. In this letter, we evaluate VENOM's performance against baselines and show that we can achieve a stable motion tracker across different humanoids more capable than an MLP trained on multiple humanoid data with supervised learning alone, and also show that despite lack of reward feedback, VENOM closely matches the tracking capability of experts that were trained using asymmetric-actor critic reinforcement learning.
翻译:从演示数据中实现跨多种类人机器人的专家级全身运动追踪,仍是类人机器人学习中尚未充分探索的挑战性问题。现有跨实体运动追踪策略主要将控制问题分解为上半身与下半身独立控制。本文提出VENOM,一种面向仿真环境下类人机器人的跨实体全身运动追踪模型。作为基于GPT架构的运动追踪器,VENOM在多种类人机器人数据上训练,无需分体控制即可实现全身追踪。我们构建了包含状态、动作与奖励的多类人机器人运动追踪数据集VENOM数据集,并基于该数据集训练VENOM及基线模型。实验表明:相比仅通过监督学习在多种类人数据上训练的MLP模型,VENOM能实现更稳定的跨实体运动追踪;且尽管缺乏奖励反馈,其追踪能力仍接近采用非对称演员-评论家强化学习训练的专家模型。