Reward specification is a notoriously difficult problem in reinforcement learning, requiring extensive expert supervision to design robust reward functions. Imitation learning (IL) methods attempt to circumvent these problems by utilizing expert demonstrations but typically require a large number of in-domain expert demonstrations. Inspired by advances in the field of Video-and-Language Models (VLMs), we present RoboCLIP, an online imitation learning method that uses a single demonstration (overcoming the large data requirement) in the form of a video demonstration or a textual description of the task to generate rewards without manual reward function design. Additionally, RoboCLIP can also utilize out-of-domain demonstrations, like videos of humans solving the task for reward generation, circumventing the need to have the same demonstration and deployment domains. RoboCLIP utilizes pretrained VLMs without any finetuning for reward generation. Reinforcement learning agents trained with RoboCLIP rewards demonstrate 2-3 times higher zero-shot performance than competing imitation learning methods on downstream robot manipulation tasks, doing so using only one video/text demonstration.
翻译:奖励规范是强化学习中公认的难题,需要大量的专家监督来设计稳健的奖励函数。模仿学习方法试图通过利用专家演示来规避这些问题,但通常需要大量的领域内专家演示。受视频与语言模型领域进展的启发,我们提出RoboCLIP,一种在线模仿学习方法。该方法只需单个视频演示或任务文本描述形式的一次演示(克服了大数据需求),即可生成奖励,无需手动设计奖励函数。此外,RoboCLIP还能利用领域外演示(例如人类解决任务的视频)进行奖励生成,从而避免演示域与部署域必须相同的要求。RoboCLIP无需任何微调即可利用预训练的视频与语言模型生成奖励。使用RoboCLIP奖励训练的强化学习智能体,在下游机器人操作任务中的零样本性能比同类模仿学习方法高出2-3倍,且仅需单个视频或文本演示。