Imitation learning has been widely applied to various autonomous systems thanks to recent development in interactive algorithms that address covariate shift and compounding errors induced by traditional approaches like behavior cloning. However, existing interactive imitation learning methods assume access to one perfect expert. Whereas in reality, it is more likely to have multiple imperfect experts instead. In this paper, we propose MEGA-DAgger, a new DAgger variant that is suitable for interactive learning with multiple imperfect experts. First, unsafe demonstrations are filtered while aggregating the training data, so the imperfect demonstrations have little influence when training the novice policy. Next, experts are evaluated and compared on scenarios-specific metrics to resolve the conflicted labels among experts. Through experiments in autonomous racing scenarios, we demonstrate that policy learned using MEGA-DAgger can outperform both experts and policies learned using the state-of-the-art interactive imitation learning algorithms such as Human-Gated DAgger. The supplementary video can be found at \url{https://youtu.be/wPCht31MHrw}.
翻译:模仿学习已广泛应用于各类自主系统,这得益于近期交互式算法的发展——该类算法能有效解决传统方法(如行为克隆)引发的协变量偏移和累积误差问题。然而,现有交互式模仿学习方法均假设可获取单个完美专家,而现实中更可能出现的是同时存在多个非完美专家。本文提出MEGA-DAgger——一种适用于多非完美专家交互式学习的新型DAgger变体。该方法首先在聚合训练数据时过滤不安全演示,从而减少非完美演示对新手策略训练的影响;其次,通过场景特定指标评估并比较各专家,解决专家间标签冲突问题。在自主竞速场景中的实验表明,采用MEGA-DAgger学习的策略不仅优于各专家,还优于采用Human-Gated DAgger等先进交互式模仿学习算法得到的策略。补充视频见\url{https://youtu.be/wPCht31MHrw}。