We propose a theoretical framework for studying behavior cloning of complex expert demonstrations using generative modeling. Our framework invokes low-level controllers - either learned or implicit in position-command control - to stabilize imitation around expert demonstrations. We show that with (a) a suitable low-level stability guarantee and (b) a powerful enough generative model as our imitation learner, pure supervised behavior cloning can generate trajectories matching the per-time step distribution of essentially arbitrary expert trajectories in an optimal transport cost. Our analysis relies on a stochastic continuity property of the learned policy we call "total variation continuity" (TVC). We then show that TVC can be ensured with minimal degradation of accuracy by combining a popular data-augmentation regimen with a novel algorithmic trick: adding augmentation noise at execution time. We instantiate our guarantees for policies parameterized by diffusion models and prove that if the learner accurately estimates the score of the (noise-augmented) expert policy, then the distribution of imitator trajectories is close to the demonstrator distribution in a natural optimal transport distance. Our analysis constructs intricate couplings between noise-augmented trajectories, a technique that may be of independent interest. We conclude by empirically validating our algorithmic recommendations, and discussing implications for future research directions for better behavior cloning with generative modeling.
翻译:我们提出了一个用于研究利用生成式建模克隆复杂专家示教行为的理论框架。该框架引入低级控制器(包括学习得到的或隐含于位置指令控制中的)以稳定专家示教轨迹附近的模仿行为。研究表明,当具备(a)适当的低级稳定性保证和(b)足够强大的生成模型作为模仿学习器时,纯监督式行为克隆能够生成在最优传输成本下与任意专家轨迹逐时间步分布相匹配的轨迹。我们的分析依赖于一种称为"全变差连续性"(TVC)的学习策略随机连续性性质,并进一步证明通过结合主流数据增强方案与新颖算法技巧(在执行阶段加入增强噪声),可以在保证精度损失最小化的前提下确保TVC性质。针对以扩散模型参数化的策略,我们具体给出了保证条件:若学习器能准确估计(经噪声增强的)专家策略的分数函数,则模仿者轨迹分布与演示者分布之间的自然最优传输距离将保持较小。我们的分析构建了噪声增强轨迹间的复杂耦合关系,该技术可能具有独立研究价值。最后通过实验验证了算法建议的有效性,并探讨了利用生成式建模改进行为克隆的未来研究方向。