Visual world models have shown great potential in learning complex system dynamics. Recent advancements leverage these models as transition functions within Model Predictive Control (MPC) frameworks to solve various control tasks. When applied to robotics, however, they are limited to single-stage tasks such as reaching or grasping, and struggle with multi-stage ones that demand complex sequential planning. In this work, we introduce WorldDP, a world model framework designed for multi-stage robotic manipulation. Our hierarchical approach utilizes a high-level world model as a transition function to optimize for feasible subgoals during runtime, which are subsequently reached by a low-level Diffusion Policy. To further aid in learning dynamics and planning, we incorporate object-centric representations that decouple environmental entities and enable us to plan sequentially with respect to each. Evaluated across several robotics benchmarks, WorldDP consistently outperforms existing baselines, validating that coupling the world model's physically grounded planning with diffusion policy's efficient execution yields superior multi-stage performance.
翻译:视觉世界模型在学习复杂系统动力学方面展现出巨大潜力。近期进展将这些模型作为模型预测控制框架中的转移函数应用于各类控制任务。然而,在机器人领域,它们仅局限于单阶段任务(如抓取或触碰),难以完成需要复杂顺序规划的多阶段任务。本研究提出WorldDP——一种面向多阶段机器人操作的世界模型框架。我们的分层方法利用高层世界模型作为转移函数,在运行时优化可行子目标,随后由底层扩散策略实现这些子目标。为进一步辅助动力学学习与规划,我们引入物体中心表示,解耦环境实体,从而能够针对每个实体进行顺序规划。在多个机器人基准评估中,WorldDP持续优于现有基线模型,验证了将世界模型的物理驱动规划与扩散策略的高效执行相结合,能够显著提升多阶段任务性能。