The recent trend in scaling models for robot learning has resulted in impressive policies that can perform various manipulation tasks and generalize to novel scenarios. However, these policies continue to struggle with following instructions, likely due to the limited linguistic and action sequence diversity in existing robotics datasets. This paper introduces Task Robustness via Re-Labelling Vision-Action Robot Data (TREAD), a scalable framework that leverages large Vision-Language Models (VLMs) to augment existing robotics datasets without additional data collection, harnessing the transferable knowledge embedded in these models. Our approach leverages a pretrained VLM through three stages: generating semantic sub-tasks from original instruction labels and initial scenes, segmenting demonstration videos conditioned on these sub-tasks, and producing diverse instructions that incorporate object properties, effectively decomposing longer demonstrations into grounded language-action pairs. We further enhance robustness by augmenting the data with linguistically diverse versions of the text goals. Evaluations on LIBERO demonstrate that policies trained on our augmented datasets exhibit improved performance on novel, unseen tasks and goals. Our results show that TREAD enhances both planning generalization through trajectory decomposition and language-conditioned policy generalization through increased linguistic diversity.


翻译:近年来,机器人学习中的模型规模化趋势催生了性能卓越的策略,这些策略能够执行多种操作任务并泛化至新场景。然而,受限于现有机器人数据集中语言指令与动作序列多样性的不足,这些策略在遵循指令方面仍存在困难。本文提出"通过重标注视觉-动作机器人数据实现任务鲁棒性"(TREAD),这是一个可扩展框架,利用大型视觉语言模型(VLM)在不额外采集数据的情况下增强现有机器人数据集,充分挖掘这些模型中蕴含的可迁移知识。该方法通过三个阶段部署预训练VLM:从原始指令标签和初始场景生成语义子任务、基于子任务对演示视频进行分割、生成融入物体属性的多样化指令,从而将长时序演示有效分解为具身化的语言-动作对。同时,我们通过引入语言多样化的文本目标来增强数据鲁棒性。在LIBERO基准上的评估表明,基于增强数据集训练的策略在应对未见任务与目标时展现出更优性能。实验结果证实,TREAD既能通过轨迹分解提升规划泛化能力,又能通过语言多样性增强提升语言条件策略的泛化能力。

0
下载
关闭预览

相关内容

《关键任务型人工智能的可靠性》
专知会员服务
20+阅读 · 4月9日
生成式人工智能在机器人操作中的应用:综述
专知会员服务
30+阅读 · 2025年3月6日
【MIT博士论文】理解与提升机器学习模型的表征鲁棒性
专知会员服务
30+阅读 · 2024年8月26日
【MIT博士论文】实用机器学习的高效鲁棒算法,142页pdf
专知会员服务
60+阅读 · 2022年9月7日
基于数据的分布式鲁棒优化算法及其应用【附PPT与视频资料】
人工智能前沿讲习班
27+阅读 · 2018年12月13日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员