Identifying causal structure is central to many fields ranging from strategic decision-making to biology and economics. In this work, we propose a model-based reinforcement learning method for causal discovery based on tree search, which builds directed acyclic graphs incrementally. We also formalize and prove the correctness of an efficient algorithm for excluding edges that would introduce cycles, which enables deeper discrete search and sampling in DAG space. We evaluate our approach on two real-world tasks, achieving substantially better performance than the state-of-the-art model-free method and greedy search, constituting a promising advancement for combinatorial methods.
翻译:识别因果结构是战略决策、生物学和经济学等多个领域的核心问题。本文提出一种基于树搜索的模型强化学习方法用于因果发现,该方法逐步构建有向无环图。我们同时形式化并证明了一种高效排除可能引入环的边的算法的正确性,该算法支持在DAG空间中进行更深入的离散搜索和采样。我们在两个实际任务上评估了该方法,相较于当前最先进的无模型方法和贪婪搜索取得了显著更优的性能,为组合优化方法的发展开辟了有前景的进展。