A crucial design decision for any robot learning pipeline is the choice of policy representation: what type of model should be used to generate the next set of robot actions? Owing to the inherent multi-modal nature of many robotic tasks, combined with the recent successes in generative modeling, researchers have turned to state-of-the-art probabilistic models such as diffusion models for policy representation. In this work, we revisit the choice of energy-based models (EBM) as a policy class. We show that the prevailing folklore -- that energy models in high dimensional continuous spaces are impractical to train -- is false. We develop a practical training objective and algorithm for energy models which combines several key ingredients: (i) ranking noise contrastive estimation (R-NCE), (ii) learnable negative samplers, and (iii) non-adversarial joint training. We prove that our proposed objective function is asymptotically consistent and quantify its limiting variance. On the other hand, we show that the Implicit Behavior Cloning (IBC) objective is actually biased even at the population level, providing a mathematical explanation for the poor performance of IBC trained energy policies in several independent follow-up works. We further extend our algorithm to learn a continuous stochastic process that bridges noise and data, modeling this process with a family of EBMs indexed by scale variable. In doing so, we demonstrate that the core idea behind recent progress in generative modeling is actually compatible with EBMs. Altogether, our proposed training algorithms enable us to train energy-based models as policies which compete with -- and even outperform -- diffusion models and other state-of-the-art approaches in several challenging multi-modal benchmarks: obstacle avoidance path planning and contact-rich block pushing.
翻译:机器人学习流程中的一个关键设计决策是策略表示的选择:应该使用哪种类型的模型来生成下一组机器人动作?鉴于许多机器人任务固有的多模态特性,结合生成式建模的最新成功,研究人员已转向使用扩散模型等最先进的概率模型进行策略表示。在这项工作中,我们重新审视了将基于能量的模型作为策略类别的选择。我们表明,普遍认为的高维连续空间中能量模型难以训练的观点是错误的。我们为能量模型开发了一种实用的训练目标和算法,该算法结合了以下几个关键要素:(i)排序噪声对比估计,(ii)可学习的负采样器,以及(iii)非对抗性联合训练。我们证明所提出的目标函数是渐近一致的,并量化了其极限方差。另一方面,我们表明,即使在总体水平上,隐式行为克隆目标实际上是有偏的,这为若干独立后续工作中基于隐式行为克隆训练的能量策略性能不佳提供了数学解释。我们进一步扩展了算法,以学习一个连接噪声和数据的连续随机过程,并通过一系列由尺度变量索引的能量模型对该过程进行建模。通过这样做,我们证明了最近生成式建模进展背后的核心思想实际上与能量模型兼容。总而言之,我们提出的训练算法使我们能够训练基于能量的模型作为策略,这些策略在多个具有挑战性的多模态基准测试中——包括避障路径规划和接触丰富的块推任务——能够与扩散模型和其他最先进方法竞争,甚至超越它们。