Learning from demonstrations (LfD) has successfully trained robots to exhibit remarkable generalization capabilities. However, many powerful imitation techniques do not prioritize the feasibility of the robot behaviors they generate. In this work, we explore the feasibility of plans produced by LfD. As in prior work, we employ a temporal diffusion model with fixed start and goal states to facilitate imitation through in-painting. Unlike previous studies, we apply cold diffusion to ensure the optimization process is directed through the agent's replay buffer of previously visited states. This routing approach increases the likelihood that the final trajectories will predominantly occupy the feasible region of the robot's state space. We test this method in simulated robotic environments with obstacles and observe a significant improvement in the agent's ability to avoid these obstacles during planning.
翻译:从示范中学习已成功训练机器人展现出显著的泛化能力。然而,许多强大的模仿技术并未优先考虑其生成的机器人行为的可行性。本研究探讨了从示范中学习所产生规划的可行性。与先前研究一致,我们采用带有固定起始和目标状态的时序扩散模型,通过图像修补实现模仿。与以往研究不同,我们应用冷扩散确保优化过程经由智能体先前访问状态的重放缓冲器进行引导。这种路由方法增加了最终轨迹主要占据机器人状态空间可行区域的可能性。我们在含障碍物的模拟机器人环境中测试该方法,观察到智能体在规划过程中避免障碍的能力显著提升。