Standard imitation learning usually assumes that demonstrations are drawn from an optimal policy distribution. However, in real-world scenarios, every human demonstration may exhibit nearly random behavior and collecting high-quality human datasets can be quite costly. This requires imitation learning can learn from imperfect demonstrations to obtain robotic policies that align human intent. Prior work uses confidence scores to extract useful information from imperfect demonstrations, which relies on access to ground truth rewards or active human supervision. In this paper, we propose a dynamics-based method to evaluate the data confidence scores without above efforts. We develop a generalized confidence-based imitation learning framework called Confidence-based Inverse soft-Q Learning (CIQL), which can employ different optimal policy matching methods by simply changing object functions. Experimental results show that our confidence evaluation method can increase the success rate by $40.3\%$ over the original algorithm and $13.5\%$ over the simple noise filtering.
翻译:标准模仿学习通常假设示范来自最优策略分布。然而,在现实场景中,每个人类示范都可能表现出近乎随机的行为,且收集高质量人类数据集成本高昂。这就要求模仿学习能够从不完美示范中学习,以获取符合人类意图的机器人策略。先前工作利用置信度分数从不完美示范中提取有用信息,这依赖于对真实奖励的访问或主动的人类监督。本文提出一种无需上述工作的基于动力学的数据置信度评估方法。我们开发了一个通用的基于置信度的模仿学习框架——置信度逆软Q学习(Confidence-based Inverse soft-Q Learning, CIQL),该框架只需通过改变目标函数即可采用不同的最优策略匹配方法。实验结果表明,我们的置信度评估方法相较于原始算法可将成功率提升$40.3\%$,相较于简单噪声过滤方法提升$13.5\%$。