Inverse reinforcement learning (IRL) aims to recover the reward function of an expert agent from demonstrations of behavior. It is well known that the IRL problem is fundamentally ill-posed, i.e., many reward functions can explain the demonstrations. For this reason, IRL has been recently reframed in terms of estimating the feasible reward set, thus, postponing the selection of a single reward. However, so far, the available formulations and algorithmic solutions have been proposed and analyzed mainly for the online setting, where the learner can interact with the environment and query the expert at will. This is clearly unrealistic in most practical applications, where the availability of an offline dataset is a much more common scenario. In this paper, we introduce a novel notion of feasible reward set capturing the opportunities and limitations of the offline setting and we analyze the complexity of its estimation. This requires the introduction an original learning framework that copes with the intrinsic difficulty of the setting, for which the data coverage is not under control. Then, we propose two computationally and statistically efficient algorithms, IRLO and PIRLO, for addressing the problem. In particular, the latter adopts a specific form of pessimism to enforce the novel desirable property of inclusion monotonicity of the delivered feasible set. With this work, we aim to provide a panorama of the challenges of the offline IRL problem and how they can be fruitfully addressed.
翻译:逆强化学习旨在从行为示范中恢复专家智能体的奖励函数。众所周知,逆强化学习问题本质上是不适定的,即存在多个奖励函数均可解释示范行为。基于此,近期研究将逆强化学习重新定义为估计可行奖励集的问题,从而推迟单一奖励的选取。然而,现有公式化表述与算法解决方案主要针对在线设定进行分析——在该设定中,学习器可与环境交互并随时向专家查询。这在大多数实际应用中显然不现实,而离线数据集的可用性才是更常见的场景。本文提出一种捕捉离线设定机遇与局限性的新型可行奖励集概念,并分析其估计复杂度。这需要引入一个应对数据覆盖率不可控难题的原始学习框架。进而,我们提出两种兼具计算与统计效率的算法(IRLO与PIRLO)来解决该问题。其中,后者采用特定形式的悲观主义,以强制执行所传递可行集的包含单调性这一新型理想性质。本研究旨在全景式展现离线逆强化学习问题面临的挑战及其有效应对方案。