Causal confusion is a phenomenon where an agent learns a policy that reflects imperfect spurious correlations in the data. Such a policy may falsely appear to be optimal during training if most of the training data contain such spurious correlations. This phenomenon is particularly pronounced in domains such as robotics, with potentially large gaps between the open- and closed-loop performance of an agent. In such settings, causally confused models may appear to perform well according to open-loop metrics during training but fail catastrophically when deployed in the real world. In this paper, we study causal confusion in offline reinforcement learning. We investigate whether selectively sampling appropriate points from a dataset of demonstrations may enable offline reinforcement learning agents to disambiguate the underlying causal mechanisms of the environment, alleviate causal confusion in offline reinforcement learning, and produce a safer model for deployment. To answer this question, we consider a set of tailored offline reinforcement learning datasets that exhibit causal ambiguity and assess the ability of active sampling techniques to reduce causal confusion at evaluation. We provide empirical evidence that uniform and active sampling techniques are able to consistently reduce causal confusion as training progresses and that active sampling is able to do so significantly more efficiently than uniform sampling.
翻译:因果混淆是指智能体学习到的策略反映了数据中不完美的虚假相关性。若大部分训练数据包含此类虚假相关性,此类策略可能在训练期间假象地呈现最优表现。该现象在机器人等域中尤为显著——智能体的开环与闭环性能可能存在巨大差距。在此类场景下,因果混淆模型可能在训练时根据开环指标表现良好,但在实际部署时灾难性失效。本文研究离线强化学习中的因果混淆问题,探究从示范数据集中选择性采样合适数据点能否使离线强化学习智能体区分环境中的潜在因果机制,缓解离线强化学习中的因果混淆,并生成更安全的部署模型。为解答该问题,我们设计了一组存在因果模糊性的定制离线强化学习数据集,评估主动采样技术在评估阶段减少因果混淆的能力。实验证据表明:随着训练推进,均匀采样与主动采样技术均能持续减少因果混淆,且主动采样的效率显著高于均匀采样。