Learning from unstructured and uncurated data has become the dominant paradigm for generative approaches in language and vision. Such unstructured and unguided behavior data, commonly known as play, is also easier to collect in robotics but much more difficult to learn from due to its inherently multimodal, noisy, and suboptimal nature. In this paper, we study this problem of learning goal-directed skill policies from unstructured play data which is labeled with language in hindsight. Specifically, we leverage advances in diffusion models to learn a multi-task diffusion model to extract robotic skills from play data. Using a conditional denoising diffusion process in the space of states and actions, we can gracefully handle the complexity and multimodality of play data and generate diverse and interesting robot behaviors. To make diffusion models more useful for skill learning, we encourage robotic agents to acquire a vocabulary of skills by introducing discrete bottlenecks into the conditional behavior generation process. In our experiments, we demonstrate the effectiveness of our approach across a wide variety of environments in both simulation and the real world. Results visualizations and videos at https://play-fusion.github.io
翻译:从非结构化和未整理数据中学习已成为语言与视觉领域生成式方法的主流范式。这类非结构化、无引导的行为数据(通常称为“游戏数据”)在机器人领域更易收集,但由于其固有的多模态、噪声和次优特性,从中学习面临更大挑战。本文研究如何从带有事后语言标签的非结构化游戏数据中学习目标导向的技能策略。具体而言,我们利用扩散模型的进展,学习一个多任务扩散模型以从游戏数据中提取机器人技能。通过在状态-动作空间中使用条件去噪扩散过程,我们能优雅地处理游戏数据的复杂性和多模态性,并生成多样且有趣的机器人行为。为提升扩散模型在技能学习中的实用性,我们通过在条件行为生成过程中引入离散瓶颈,促使机器人智能体习得技能词汇表。实验表明,本方法在仿真和现实世界的多种环境中均展现出有效性。结果可视化及视频见 https://play-fusion.github.io