We consider offline imitation learning (IL), which aims to mimic the expert's behavior from its demonstration without further interaction with the environment. One of the main challenges in offline IL is dealing with the limited support of expert demonstrations that cover only a small fraction of the state-action spaces. In this work, we consider offline IL, where expert demonstrations are limited but complemented by a larger set of sub-optimal demonstrations of lower expertise levels. Most of the existing offline IL methods developed for this setting are based on behavior cloning or distribution matching, where the aim is to match the occupancy distribution of the imitation policy with that of the expert policy. Such an approach often suffers from over-fitting, as expert demonstrations are limited to accurately represent any occupancy distribution. On the other hand, since sub-optimal sets are much larger, there is a high chance that the imitation policy is trained towards sub-optimal policies. In this paper, to address these issues, we propose a new approach based on inverse soft-Q learning, where a regularization term is added to the training objective, with the aim of aligning the learned rewards with a pre-assigned reward function that allocates higher weights to state-action pairs from expert demonstrations, and lower weights to those from lower expertise levels. On standard benchmarks, our inverse soft-Q learning significantly outperforms other offline IL baselines by a large margin.
翻译:摘要:本文研究离线模仿学习(IL),旨在无需与环境进一步交互的情况下,通过专家示范模仿其行为。离线IL的主要挑战之一是处理仅覆盖状态-动作空间极小部分的专家示范的有限支持问题。本研究考虑专家示范有限但辅以更多低专业水平的次优示范集的离线IL场景。现有针对该场景的大多数离线IL方法基于行为克隆或分布匹配,目标是将模仿策略的占据分布与专家策略的占据分布对齐。此类方法常因专家示范不足而难以准确表征任意占据分布,导致过拟合。另一方面,由于次优示范集规模更大,模仿策略极有可能趋向于学习次优策略。为解决这些问题,本文提出一种基于逆向软Q学习的新方法,在训练目标中添加正则化项,使学习到的奖励函数与预设奖励函数对齐——该预设函数对来自专家示范的状态动作对赋予更高权重,对来自低专业水平的示范赋予更低权重。在标准基准测试中,我们的逆向软Q学习方法显著优于其他离线IL基线方法,性能大幅领先。