Offline reinforcement learning (RL) aims to learn a policy using only pre-collected and fixed data. Although avoiding the time-consuming online interactions in RL, it poses challenges for out-of-distribution (OOD) state actions and often suffers from data inefficiency for training. Despite many efforts being devoted to addressing OOD state actions, the latter (data inefficiency) receives little attention in offline RL. To address this, this paper proposes the cross-domain offline RL, which assumes offline data incorporate additional source-domain data from varying transition dynamics (environments), and expects it to contribute to the offline data efficiency. To do so, we identify a new challenge of OOD transition dynamics, beyond the common OOD state actions issue, when utilizing cross-domain offline data. Then, we propose our method BOSA, which employs two support-constrained objectives to address the above OOD issues. Through extensive experiments in the cross-domain offline RL setting, we demonstrate BOSA can greatly improve offline data efficiency: using only 10\% of the target data, BOSA could achieve {74.4\%} of the SOTA offline RL performance that uses 100\% of the target data. Additionally, we also show BOSA can be effortlessly plugged into model-based offline RL and noising data augmentation techniques (used for generating source-domain data), which naturally avoids the potential dynamics mismatch between target-domain data and newly generated source-domain data.
翻译:离线强化学习旨在仅使用预先收集的固定数据学习策略。尽管避免了强化学习中耗时的在线交互,但它面临分布外状态动作的挑战,且常因训练数据效率低下而受限。尽管许多研究致力于解决OOD状态动作问题,但数据效率低下在离线RL中鲜受关注。为此,本文提出跨域离线RL方法,假设离线数据包含来自不同转移动态(环境)的额外源域数据,并期望其提升离线数据效率。针对跨域离线数据的应用,我们识别出超越常见OOD状态动作问题的新挑战——OOD转移动态。随后提出BOSA方法,采用两个支持约束目标来解决上述OOD问题。通过在跨域离线RL设置下的广泛实验,我们证明BOSA能显著提升离线数据效率:仅使用10%的目标域数据,BOSA即可达到使用100%目标域数据的SOTA离线RL性能的74.4%。此外,我们进一步展示BOSA可无缝嵌入基于模型的离线RL和噪声数据增强技术(用于生成源域数据),从而自然规避目标域数据与新生成源域数据之间潜在的动态不匹配问题。