Causal discovery aims to uncover causal structures from observational data, which is crucial for real-world decision-making. However, different causal discovery algorithms can produce divergent results that conflict with each other, complicating the identification of accurate causal graphs. Traditional approaches rely on numerical values and statistical assumptions, often ignoring rich domain-specific information, such as feature descriptions, which could also help structure learning. While recent works explore using Large Language Models (LLMs) to infer causal relations via direct queries, such methods can be unreliable due to a lack of alignment with the actual data. To address these limitations, we propose Causal Ensemble Agent (CEA), a novel framework that aggregates structural insights from statistical discovery experts across different graph levels via linear opinion pooling, and uses an LLM as a meta-referee to dynamically reweight experts when the aggregated confidence is close to the decision boundary, thereby composing an improved and more complete causal graph. Extensive experiments on both synthetic and real-world datasets demonstrate that CEA achieves the strongest overall performance across a wide range of causal discovery methods, highlighting the effectiveness of using LLMs for meta-analysis in causal discovery.
翻译:因果发现旨在从观测数据中揭示因果结构,这对现实世界的决策至关重要。然而,不同的因果发现算法可能产生相互矛盾的结果,使准确因果图的识别变得复杂。传统方法依赖于数值和统计假设,往往忽略了丰富的领域特定信息(如特征描述),而这些信息本可辅助结构学习。尽管近期研究尝试利用大型语言模型(LLMs)通过直接查询推断因果关系,但由于缺乏与实际数据的一致性,此类方法可能并不可靠。为解决这些局限,我们提出因果集成智能体(CEA),该新型框架通过线性意见池化,聚合来自不同图层次的统计发现专家的结构见解,并在聚合置信度接近决策边界时,利用LLM作为元裁判动态重新加权专家,从而组合出更优且更完整的因果图。在合成数据集和真实数据集上的大量实验表明,CEA在广泛的因果发现方法中实现了最强的整体性能,凸显了LLM在因果发现元分析中的有效性。