A prevailing view in robot learning is that simulation alone is not enough; effective sim-to-real transfer is widely believed to require at least some real-world data collection or task-specific fine-tuning to bridge the gap between simulated and physical environments. We challenge that assumption. With sufficiently large-scale and diverse simulated synthetic training data, we show that zero-shot transfer to the real world is not only possible, but effective for both static and mobile manipulation. We introduce MolmoBot-Engine, a fully open-source pipeline for procedural data generation across robots, tasks, and diverse simulated environments in MolmoSpaces. With it, we release MolmoBot-Data, a dataset of 1.8 million expert trajectories for articulated object manipulation and pick-and-place tasks. We train three policy classes: MolmoBot, a Molmo2-based multi-frame vision-language model with a flow-matching action head; MolmoBot-Pi0, which replicates the $π_0$ architecture to enable direct comparison; and MolmoBot-SPOC, a lightweight policy suitable for edge deployment and amenable to RL fine-tuning. We evaluate on two robotic platforms: the Franka FR3 for tabletop manipulation tasks and the Rainbow Robotics RB-Y1 mobile manipulator for door opening, drawer manipulation, cabinet interaction, and mobile pick-and-place. Without any real-world fine-tuning, our policies achieve zero-shot transfer to unseen objects and environments. On tabletop pick-and-place, MolmoBot achieves a success rate of 79.2% in real world evaluations across 4 settings, outperforming $π_{0.5}$ at 39.2%. Our results demonstrate that procedural environment generation combined with diverse articulated assets can produce robust manipulation policies that generalize broadly to the real world. Technical website: https://allenai.github.io/MolmoBot


翻译:机器人学习领域的主流观点认为,仅靠模拟是不够的;有效的模拟到现实迁移通常需要至少一部分真实世界数据收集或任务特定微调,以弥合模拟环境与物理环境之间的差距。我们挑战这一假设。研究表明,通过足够大规模且多样化的模拟合成训练数据,零样本迁移至真实世界不仅在静态操作中可行,在移动操作中也同样有效。我们提出MolmoBot-Engine,一个完全开源的流水线,用于在MolmoSpaces中跨机器人、任务和多样化模拟环境进行程序化数据生成。依托该引擎,我们发布了MolmoBot-Data数据集,包含180万条针对铰接物体操作和抓取放置任务的专家轨迹。我们训练了三种策略类别:MolmoBot——一种基于Molmo2的多帧视觉语言模型,配备流匹配动作头;MolmoBot-Pi0——复现$π_0$架构以支持直接对比;以及MolmoBot-SPOC——一种适用于边缘部署且便于强化学习微调的轻量级策略。我们在两个机器人平台上进行评估:用于桌面操作任务的Franka FR3,以及用于开门、抽屉操作、柜体交互和移动抓取放置的Rainbow Robotics RB-Y1移动操作器。无需任何真实世界微调,我们的策略即可在未见物体和环境中实现零样本迁移。在桌面抓取放置任务中,MolmoBot在4种场景下的真实世界评估成功率可达79.2%,优于$π_{0.5}$的39.2%。结果表明,程序化环境生成与多样化铰接资产相结合,能够产生可广泛泛化至真实世界的鲁棒操作策略。技术网站:https://allenai.github.io/MolmoBot

0
下载
关闭预览

相关内容

【自动化学报】零样本学习研究进展,中国石油大学
专知会员服务
88+阅读 · 2020年1月27日
零样本图像识别综述论文
专知
22+阅读 · 2020年4月4日
入门 | 深度学习模型的简单优化技巧
机器之心
10+阅读 · 2018年6月10日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2013年12月31日
VIP会员
最新内容
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
1+阅读 · 今天13:43
博士论文 | 大动作空间中的在线与离线策略学习
专知会员服务
0+阅读 · 今天13:36
综述 | Autonomous Research Agents:AI 科学家与验证缺口
《多域冲突比较支持模型》60页
专知会员服务
9+阅读 · 8月7日
面向2027年及未来的海军情报改革
专知会员服务
6+阅读 · 8月5日
相关VIP内容
【自动化学报】零样本学习研究进展,中国石油大学
专知会员服务
88+阅读 · 2020年1月27日
相关基金
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2013年12月31日
Top
微信扫码咨询专知VIP会员