As a pivotal component to attaining generalizable solutions in human intelligence, reasoning provides great potential for reinforcement learning (RL) agents' generalization towards varied goals by summarizing part-to-whole arguments and discovering cause-and-effect relations. However, how to discover and represent causalities remains a huge gap that hinders the development of causal RL. In this paper, we augment Goal-Conditioned RL (GCRL) with Causal Graph (CG), a structure built upon the relation between objects and events. We novelly formulate the GCRL problem into variational likelihood maximization with CG as latent variables. To optimize the derived objective, we propose a framework with theoretical performance guarantees that alternates between two steps: using interventional data to estimate the posterior of CG; using CG to learn generalizable models and interpretable policies. Due to the lack of public benchmarks that verify generalization capability under reasoning, we design nine tasks and then empirically show the effectiveness of the proposed method against five baselines on these tasks. Further theoretical analysis shows that our performance improvement is attributed to the virtuous cycle of causal discovery, transition modeling, and policy training, which aligns with the experimental evidence in extensive ablation studies.
翻译:作为人类智能中实现可泛化解的关键组件,推理通过归纳整体与部分的论证关系并发现因果联系,为强化学习智能体实现面向多样化目标的泛化提供了巨大潜力。然而,如何发现并表征因果关系仍是阻碍因果强化学习发展的重大难题。本文通过构建基于物体与事件关系的因果图,对目标条件强化学习进行增强。我们创新性地将目标条件强化学习问题转化为以因果图为隐变量的变分似然最大化问题。为优化导出的目标函数,我们提出一个具有理论性能保证的框架,通过交替执行两个步骤:利用干预数据估计因果图的后验分布;利用因果图学习可泛化模型与可解释策略。针对当前缺乏验证推理泛化能力的公开基准,我们设计九项任务,并在这些任务上通过实证验证所提方法相较于五种基准方法的有效性。进一步的理论分析表明,我们的性能提升归因于因果发现、转移建模与策略训练形成的良性循环,这与广泛消融实验中的证据高度吻合。