Embodied control requires agents to leverage multi-modal pre-training to quickly learn how to act in new environments, where video demonstrations contain visual and motion details needed for low-level perception and control, and language instructions support generalization with abstract, symbolic structures. While recent approaches apply contrastive learning to force alignment between the two modalities, we hypothesize better modeling their complementary differences can lead to more holistic representations for downstream adaption. To this end, we propose Emergent Communication for Embodied Control (EC^2), a novel scheme to pre-train video-language representations for few-shot embodied control. The key idea is to learn an unsupervised "language" of videos via emergent communication, which bridges the semantics of video details and structures of natural language. We learn embodied representations of video trajectories, emergent language, and natural language using a language model, which is then used to finetune a lightweight policy network for downstream control. Through extensive experiments in Metaworld and Franka Kitchen embodied benchmarks, EC^2 is shown to consistently outperform previous contrastive learning methods for both videos and texts as task inputs. Further ablations confirm the importance of the emergent language, which is beneficial for both video and language learning, and significantly superior to using pre-trained video captions. We also present a quantitative and qualitative analysis of the emergent language and discuss future directions toward better understanding and leveraging emergent communication in embodied tasks.
翻译:具身控制要求智能体利用多模态预训练快速学习在新环境中行动,其中视频演示包含低层感知与控制所需的视觉和运动细节,而语言指令则通过抽象的符号结构支持泛化。尽管近期方法采用对比学习强制两种模态之间的对齐,但我们假设更好地建模它们的互补差异可以产生更全面的表征以支持下游适应。为此,我们提出用于具身控制的涌现通信(EC^2),这是一种新颖的视频-语言表征预训练方案,旨在实现小样本具身控制。其核心思想是通过涌现通信学习视频的无监督“语言”,该语言桥接了视频细节的语义与自然语言的结构。我们利用语言模型学习视频轨迹、涌现语言和自然语言的表征,随后使用该模型微调轻量级策略网络以进行下游控制。通过在Metaworld和Franka Kitchen具身基准上的大量实验,EC^2被证明在视频和文本作为任务输入时始终优于以往的对比学习方法。进一步的消融实验证实了涌现语言的重要性,它同时有益于视频和语言学习,且显著优于使用预训练的视频字幕。我们还对涌现语言进行了定量和定性分析,并讨论了在具身任务中更好理解和利用涌现通信的未来方向。