Achieving versatile humanoid locomotion with a single policy presents a critical scalability challenge. Prevailing methods often rely on distilling multiple terrain-specific teacher policies into a unified student policy. However, while such distillation captures basic locomotion primitives, it struggles to organically compose these skills to adapt to complex environments, resulting in poor generalization to novel composite terrains unseen during training. To overcome this, we present DreamPolicy, a unified framework that integrates offline data with a diffusion-based world model, enabling a single policy to master both known and unseen terrains. Central to our approach is a terrain-aware world model, driven by an autoregressive diffusion world model trained on aggregated rollouts from specialized policies. This model synthesizes physically plausible future trajectories, which serve as dynamic objectives for a conditioned policy, thereby bypassing manual reward engineering. Unlike distillation, our world model captures generalizable locomotion skills, allowing for robust zero-shot transfer to unseen composite terrains. DreamPolicy naturally scales with data availability. As the offline dataset expands, the diffusion world model continuously acquires richer skills. Experiments demonstrate that DreamPolicy outperforms the strongest baseline by up to 27\% on unseen terrains and 38\% on combined terrains. By unifying world model-based planning and policy learning, DreamPolicy breaks the "one task, one policy" bottleneck and establishes a scalable, data-driven paradigm for generalist humanoid control.


翻译:实现单一策略驱动的通用人形运动面临关键的可扩展性挑战。现有方法通常依赖将多个特定地形教师策略蒸馏为统一的学生策略,但此类蒸馏虽能捕捉基础运动基元,却难以有机组合这些技能以适应复杂环境,导致对训练中未见的新型组合地形泛化能力差。为攻克这一难题,我们提出DreamPolicy——将离线数据与基于扩散的世界模型相融合的统一框架,使单一策略既能掌握已知地形,也能适应未知地形。该框架的核心是地形感知世界模型,该模型由基于自回归扩散的世界模型驱动,并在专用策略生成的聚合轨迹上训练。该模型可合成物理上合理的未来轨迹,作为条件策略的动态目标,从而绕过人工奖励工程。与蒸馏不同,我们的世界模型可捕获具有泛化能力的运动技能,实现对未见组合地形的鲁棒零样本迁移。DreamPolicy天然具备随数据规模扩展的能力:随着离线数据集扩大,扩散世界模型持续习得更丰富的技能。实验表明,DreamPolicy在未见地形上性能超越最强基线达27%,在组合地形上提升达38%。通过统一基于世界模型的规划与策略学习,DreamPolicy打破了"单任务单策略"的瓶颈,为通用人形控制建立了可扩展、数据驱动的范式。

0
下载
关闭预览

相关内容

《异构人类团队的协作决策过程混合建模研究》
专知会员服务
10+阅读 · 7月28日
【综述】 机器人学习中的世界模型:全面综述
专知会员服务
21+阅读 · 5月4日
走向通用人工智能之路,世界模型为何不可或缺?
专知会员服务
21+阅读 · 2025年7月1日
AI大模型驱动的具身智能人形机器人技术与展望
专知会员服务
27+阅读 · 2025年5月26日
PlaNet 简介:用于强化学习的深度规划网络
谷歌开发者
13+阅读 · 2019年3月16日
DARPA征集无人集群战术思路
无人机
19+阅读 · 2017年10月18日
国家自然科学基金
15+阅读 · 2016年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
VIP会员
最新内容
驱动军事决策变革的顶尖人工智能指挥系统
专知会员服务
6+阅读 · 8月11日
非对称防御中的自组织临界性:俄乌战争
专知会员服务
10+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
12+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
12+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
13+阅读 · 8月8日
相关基金
国家自然科学基金
15+阅读 · 2016年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员