Sequence modeling approaches have shown promising results in robot imitation learning. Recently, diffusion models have been adopted for behavioral cloning, benefiting from their exceptional capabilities in modeling complex data distribution. In this work, we propose Crossway Diffusion, a method to enhance diffusion-based visuomotor policy learning by using an extra self-supervised learning (SSL) objective. The standard diffusion-based policy generates action sequences from random noise conditioned on visual observations and other low-dimensional states. We further extend this by introducing a new decoder that reconstructs raw image pixels (and other state information) from the intermediate representations of the reverse diffusion process, and train the model jointly using the SSL loss. Our experiments demonstrate the effectiveness of Crossway Diffusion in various simulated and real-world robot tasks, confirming its advantages over the standard diffusion-based policy. We demonstrate that such self-supervised reconstruction enables better representation for policy learning, especially when the demonstrations have different proficiencies.
翻译:序列建模方法在机器人模仿学习中已展现出良好效果。近年来,扩散模型凭借其建模复杂数据分布的卓越能力被广泛应用于行为克隆。本文提出Crossway Diffusion方法,通过引入额外的自监督学习目标来增强基于扩散的视觉运动策略学习。标准扩散策略以视觉观测及其他低维状态为条件,从随机噪声生成动作序列。我们在此基础上进一步扩展,引入一个新解码器,从逆扩散过程的中间表示重建原始图像像素(及其他状态信息),并联合自监督损失训练模型。实验结果表明,Crossway Diffusion在多种仿真及真实机器人任务中具有有效性,证实其相比标准扩散策略的优势。我们证明,这种自监督重建能够为策略学习提供更优的表征,尤其当演示数据具有不同熟练度水平时效果更为显著。