Offline imitation from observations aims to solve MDPs where only task-specific expert states and task-agnostic non-expert state-action pairs are available. Offline imitation is useful in real-world scenarios where arbitrary interactions are costly and expert actions are unavailable. The state-of-the-art "DIstribution Correction Estimation" (DICE) methods minimize divergence of state occupancy between expert and learner policies and retrieve a policy with weighted behavior cloning; however, their results are unstable when learning from incomplete trajectories, due to a non-robust optimization in the dual domain. To address the issue, in this paper, we propose Trajectory-Aware Imitation Learning from Observations (TAILO). TAILO uses a discounted sum along the future trajectory as the weight for weighted behavior cloning. The terms for the sum are scaled by the output of a discriminator, which aims to identify expert states. Despite simplicity, TAILO works well if there exist trajectories or segments of expert behavior in the task-agnostic data, a common assumption in prior work. In experiments across multiple testbeds, we find TAILO to be more robust and effective, particularly with incomplete trajectories.
翻译:离线观测模仿旨在解决马尔可夫决策过程(MDP),其中仅包含任务特定的专家状态和任务无关的非专家状态-动作对。在实际场景中,当任意交互成本高昂且专家动作不可用时,离线模仿非常有用。最先进的“分布校正估计”(DICE)方法最小化专家策略与学习策略之间状态占用率的散度,并通过加权行为克隆检索策略;然而,由于对偶域中的非鲁棒优化,这些方法在从不完整轨迹学习时结果不稳定。为解决此问题,本文提出了一种基于轨迹的观测模仿学习(TAILO)。TAILO使用沿未来轨迹的折扣和作为加权行为克隆的权重。该和的各项通过判别器的输出进行缩放,该判别器旨在识别专家状态。尽管方法简单,但如果任务无关数据中存在专家行为的轨迹或片段(这是先前工作中的常见假设),TAILO仍能良好运行。在跨多个测试平台的实验中,我们发现TAILO更加鲁棒和有效,尤其是在处理不完整轨迹时。