Imitation Learning (IL) aims to discover a policy by minimizing the discrepancy between the agent's behavior and expert demonstrations. However, IL is susceptible to limitations imposed by noisy demonstrations from non-expert behaviors, presenting a significant challenge due to the lack of supplementary information to assess their expertise. In this paper, we introduce Self-Motivated Imitation LEarning (SMILE), a method capable of progressively filtering out demonstrations collected by policies deemed inferior to the current policy, eliminating the need for additional information. We utilize the forward and reverse processes of Diffusion Models to emulate the shift in demonstration expertise from low to high and vice versa, thereby extracting the noise information that diffuses expertise. Then, the noise information is leveraged to predict the diffusion steps between the current policy and demonstrators, which we theoretically demonstrate its equivalence to their expertise gap. We further explain in detail how the predicted diffusion steps are applied to filter out noisy demonstrations in a self-motivated manner and provide its theoretical grounds. Through empirical evaluations on MuJoCo tasks, we demonstrate that our method is proficient in learning the expert policy amidst noisy demonstrations, and effectively filters out demonstrations with expertise inferior to the current policy.
翻译:模仿学习旨在通过最小化智能体行为与专家演示之间的差异来发现最优策略。然而,由于缺乏评估非专家行为专业程度的辅助信息,模仿学习容易受到非专家行为产生的噪声演示的限制。本文提出了一种名为自激励模仿学习(SMILE)的方法,该方法能够逐步过滤掉由低于当前策略水平策略收集的演示,而无需额外信息。我们利用扩散模型的前向与反向过程模拟演示专业程度从低到高及相反的演变,从而提取扩散专业程度的噪声信息。随后,该噪声信息被用于预测当前策略与演示者之间的扩散步长,理论上我们证明了该步长等价于两者间的专业程度差距。我们进一步详细阐述了如何以自激励方式将预测的扩散步长应用于过滤噪声演示,并提供了其理论基础。通过在MuJoCo任务上的实证评估,我们证明了该方法能在噪声演示中有效学习专家策略,并成功过滤掉专业程度低于当前策略的演示。