The recent trend in scaling models for robot learning has resulted in impressive policies that can perform various manipulation tasks and generalize to novel scenarios. However, these policies continue to struggle with following instructions, likely due to the limited linguistic and action sequence diversity in existing robotics datasets. This paper introduces Task Robustness via Re-Labelling Vision-Action Robot Data (TREAD), a scalable framework that leverages large Vision-Language Models (VLMs) to augment existing robotics datasets without additional data collection, harnessing the transferable knowledge embedded in these models. Our approach leverages a pretrained VLM through three stages: generating semantic sub-tasks from original instruction labels and initial scenes, segmenting demonstration videos conditioned on these sub-tasks, and producing diverse instructions that incorporate object properties, effectively decomposing longer demonstrations into grounded language-action pairs. We further enhance robustness by augmenting the data with linguistically diverse versions of the text goals. Evaluations on LIBERO demonstrate that policies trained on our augmented datasets exhibit improved performance on novel, unseen tasks and goals. Our results show that TREAD enhances both planning generalization through trajectory decomposition and language-conditioned policy generalization through increased linguistic diversity.
翻译:近年来,机器人学习中的模型规模化趋势催生了性能卓越的策略,这些策略能够执行多种操作任务并泛化至新场景。然而,受限于现有机器人数据集中语言指令与动作序列多样性的不足,这些策略在遵循指令方面仍存在困难。本文提出"通过重标注视觉-动作机器人数据实现任务鲁棒性"(TREAD),这是一个可扩展框架,利用大型视觉语言模型(VLM)在不额外采集数据的情况下增强现有机器人数据集,充分挖掘这些模型中蕴含的可迁移知识。该方法通过三个阶段部署预训练VLM:从原始指令标签和初始场景生成语义子任务、基于子任务对演示视频进行分割、生成融入物体属性的多样化指令,从而将长时序演示有效分解为具身化的语言-动作对。同时,我们通过引入语言多样化的文本目标来增强数据鲁棒性。在LIBERO基准上的评估表明,基于增强数据集训练的策略在应对未见任务与目标时展现出更优性能。实验结果证实,TREAD既能通过轨迹分解提升规划泛化能力,又能通过语言多样性增强提升语言条件策略的泛化能力。