Recent works have shown that Large Language Models (LLMs) can be applied to ground natural language to a wide variety of robot skills. However, in practice, learning multi-task, language-conditioned robotic skills typically requires large-scale data collection and frequent human intervention to reset the environment or help correcting the current policies. In this work, we propose a novel approach to efficiently learn general-purpose language-conditioned robot skills from unstructured, offline and reset-free data in the real world by exploiting a self-supervised visuo-lingual affordance model, which requires annotating as little as 1% of the total data with language. We evaluate our method in extensive experiments both in simulated and real-world robotic tasks, achieving state-of-the-art performance on the challenging CALVIN benchmark and learning over 25 distinct visuomotor manipulation tasks with a single policy in the real world. We find that when paired with LLMs to break down abstract natural language instructions into subgoals via few-shot prompting, our method is capable of completing long-horizon, multi-tier tasks in the real world, while requiring an order of magnitude less data than previous approaches. Code and videos are available at http://hulc2.cs.uni-freiburg.de
翻译:近期研究表明,大型语言模型(LLMs)可应用于将自然语言与多种机器人技能进行基础对齐。然而在实践中,学习多任务、语言条件化的机器人技能通常需要大规模数据采集和频繁的人工干预(如重置环境或协助修正当前策略)。本文提出一种新方法,通过利用自监督的视觉-语言可供性模型,仅需对总数据中1%的内容进行语言标注,即可从真实世界中非结构化、离线且无需重置的数据中高效学习通用语言条件化机器人技能。我们在模拟环境和真实机器人任务中进行了广泛实验评估,在具有挑战性的CALVIN基准测试中达到当前最优性能,并在真实世界中通过单一策略学习了25种以上的不同视觉运动操控任务。研究发现,当结合LLMs通过少样本提示将抽象自然语言指令分解为子目标时,该方法能够在真实世界中完成长时域、多层级任务,同时所需数据量比先前方法少一个数量级。代码与视频见http://hulc2.cs.uni-freiburg.de