Offline reinforcement learning (RL) aims to infer sequential decision policies using only offline datasets. This is a particularly difficult setup, especially when learning to achieve multiple different goals or outcomes under a given scenario with only sparse rewards. For offline learning of goal-conditioned policies via supervised learning, previous work has shown that an advantage weighted log-likelihood loss guarantees monotonic policy improvement. In this work we argue that, despite its benefits, this approach is still insufficient to fully address the distribution shift and multi-modality problems. The latter is particularly severe in long-horizon tasks where finding a unique and optimal policy that goes from a state to the desired goal is challenging as there may be multiple and potentially conflicting solutions. To tackle these challenges, we propose a complementary advantage-based weighting scheme that introduces an additional source of inductive bias: given a value-based partitioning of the state space, the contribution of actions expected to lead to target regions that are easier to reach, compared to the final goal, is further increased. Empirically, we demonstrate that the proposed approach, Dual-Advantage Weighted Offline Goal-conditioned RL (DAWOG), outperforms several competing offline algorithms in commonly used benchmarks. Analytically, we offer a guarantee that the learnt policy is never worse than the underlying behaviour policy.
翻译:离线强化学习(RL)旨在仅利用离线数据集推断序贯决策策略。这是一个极具挑战性的设置,尤其是在仅含稀疏奖励的场景下学习实现多个不同目标或结果时。对于通过监督学习进行目标条件策略的离线学习,先前研究表明,优势加权对数似然损失可保证策略单调改进。本文指出,尽管该方案具有优势,但仍不足以完全解决分布偏移与多模态问题。后者在长时域任务中尤为严峻——当存在多种潜在冲突的解决方案时,从某个状态到期望目标寻找唯一最优策略具有挑战性。为应对这些挑战,我们提出一种互补的优势加权方案,引入额外的归纳偏置:基于状态空间的价值划分,与最终目标相比,预期导向更易到达目标区域的行动贡献将被进一步强化。实验表明,所提方法——双重优势加权离线目标条件强化学习(DAWOG)——在常用基准测试中优于多种竞争性离线算法。理论分析保证,学习到的策略性能永远不会低于底层行为策略。