Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic interfaces, but neither offers a scalable and interactive prior for broad full-body behavior. We introduce EgoPriMo (Egocentric Motion Prior for Humanoid Robots), a unified framework that learns such priors from egocentric human demonstrations. Given egocentric observations and a text prompt, EgoPriMo reconstructs, generates, and forecasts SMPL-based full-body motion. Language is used as a high-level control signal rather than a complete motion specification. At the core of EgoPriMo is a Triple-stream DiT that jointly models body dynamics, egocentric visual context, and text; task-conditioning masks route different tasks and missing-modality data through the same checkpoint. Experiments on Nymeria and EgoExo4D show that one checkpoint improves egocentric motion generation over UniEgoMotion while supporting reconstruction and forecasting; the generated SMPL motions can also be executed by a Unitree humanoid controller. These results indicate a practical path from scalable egocentric observations to generalizable and interactive humanoid motion priors.
翻译:人形机器人需要适应场景上下文、任务需求和用户意图的全身运动。运动追踪可以复现指定轨迹,而人形视觉-语言-动作系统提供了语义接口,但二者均未能为广泛的全身行为提供可扩展且交互式的先验知识。我们提出EgoPriMo(面向人形机器人的自我中心运动先验),这是一个统一框架,能从自我中心人体演示中学习此类先验。给定自我中心观察和文本提示,EgoPriMo可重建、生成并预测基于SMPL的全身运动,其中语言被用作高层控制信号而非完整的运动规范。EgoPriMo的核心是三重流DiT,它联合建模身体动力学、自我中心视觉上下文和文本;任务条件掩码通过同一检查点路由不同任务和缺失模态数据。在Nymeria和EgoExo4D上的实验表明,单一检查点在自我中心运动生成上优于UniEgoMotion,同时支持重建和预测;生成的SMPL运动还可由宇树科技人形控制器执行。这些结果揭示了一条从可扩展的自我中心观察到通用且可交互的人形运动先验的实践路径。