We study a new paradigm for sequential decision making, called offline Policy Learning from Observation (PLfO). Offline PLfO aims to learn policies using datasets with substandard qualities: 1) only a subset of trajectories is labeled with rewards, 2) labeled trajectories may not contain actions, 3) labeled trajectories may not be of high quality, and 4) the overall data may not have full coverage. Such imperfection is common in real-world learning scenarios, so offline PLfO encompasses many existing offline learning setups, including offline imitation learning (IL), ILfO, and reinforcement learning (RL). In this work, we present a generic approach, called Modality-agnostic Adversarial Hypothesis Adaptation for Learning from Observations (MAHALO), for offline PLfO. Built upon the pessimism concept in offline RL, MAHALO optimizes the policy using a performance lower bound that accounts for uncertainty due to the dataset's insufficient converge. We implement this idea by adversarially training data-consistent critic and reward functions in policy optimization, which forces the learned policy to be robust to the data deficiency. We show that MAHALO consistently outperforms or matches specialized algorithms across a variety of offline PLfO tasks in theory and experiments.
翻译:我们研究了一种新的序列决策范式,称为基于离线数据的观察驱动策略学习(Offline PLfO)。离线PLfO旨在利用存在质量缺陷的数据集学习策略,这些缺陷包括:1)仅部分轨迹带有奖励标签,2)带标签的轨迹可能不包含动作信息,3)带标签的轨迹可能并非高质量数据,4)整体数据可能缺乏完整覆盖。这种不完美性在真实学习场景中普遍存在,因此离线PLfO涵盖了许多现有的离线学习框架,包括离线模仿学习(IL)、观察驱动的模仿学习(ILfO)以及强化学习(RL)。本文提出了一种通用方法,称为模态无关的对抗性假设自适应观察学习(MAHALO),用于解决离线PLfO问题。基于离线RL中的悲观主义原则,MAHALO通过考虑数据集覆盖不足导致的不确定性,利用性能下界优化策略。我们通过在策略优化中对抗性训练数据一致的评判器与奖励函数来实现这一思想,迫使学到的策略对数据缺陷具有鲁棒性。理论和实验表明,在各种离线PLfO任务中,MAHALO在性能上始终优于或匹配专业化算法。