Reasoning about the effect of interventions and counterfactuals is a fundamental task found throughout the data sciences. A collection of principles, algorithms, and tools has been developed for performing such tasks in the last decades (Pearl, 2000). One of the pervasive requirements found throughout this literature is the articulation of assumptions, which commonly appear in the form of causal diagrams. Despite the power of this approach, there are significant settings where the knowledge necessary to specify a causal diagram over all variables is not available, particularly in complex, high-dimensional domains. In this paper, we introduce a new graphical modeling tool called cluster DAGs (for short, C-DAGs) that allows for the partial specification of relationships among variables based on limited prior knowledge, alleviating the stringent requirement of specifying a full causal diagram. A C-DAG specifies relationships between clusters of variables, while the relationships between the variables within a cluster are left unspecified, and can be seen as a graphical representation of an equivalence class of causal diagrams that share the relationships among the clusters. We develop the foundations and machinery for valid inferences over C-DAGs about the clusters of variables at each layer of Pearl's Causal Hierarchy (Pearl and Mackenzie 2018; Bareinboim et al. 2020) - L1 (probabilistic), L2 (interventional), and L3 (counterfactual). In particular, we prove the soundness and completeness of d-separation for probabilistic inference in C-DAGs. Further, we demonstrate the validity of Pearl's do-calculus rules over C-DAGs and show that the standard ID identification algorithm is sound and complete to systematically compute causal effects from observational data given a C-DAG. Finally, we show that C-DAGs are valid for performing counterfactual inferences about clusters of variables.
翻译:关于干预和反事实效应的推理是数据科学中的基础任务。过去几十年间,学术界已发展出一系列用于执行此类任务的原理、算法与工具(Pearl, 2000)。该领域文献普遍存在的基本要求是明确表述假设条件,这些假设通常以因果图的形式呈现。尽管这种方法具有强大功能,但在复杂高维领域等关键场景中,常缺乏能为所有变量指定完整因果图的必要知识。本文提出一种名为"簇有向无环图"(简称C-DAG)的新型图形建模工具,该工具允许基于有限先验知识对变量间关系进行部分规约,从而缓解了需要完整因果图的严苛要求。C-DAG在簇层面指定变量集合之间的关系,而簇内变量间关系则留待未定义,可视为等价类因果图的图形化表达——这些等价因果图共享簇间关系结构。我们建立了基于C-DAG对Pearl因果层级(Pearl and Mackenzie 2018; Bareinboim et al. 2020)各层——L1(概率层)、L2(干预层)与L3(反事实层)——进行变量簇有效推理的理论基础与机制。具体而言,我们证明了d-分离在C-DAG概率推理中的完备性与可靠性,进一步论证了Pearl do-演算规则在C-DAG上的有效性,并展示了标准ID识别算法在给定C-DAG时,可系统性地从观测数据中完备计算因果效应。最后,我们证明C-DAG可有效执行变量簇的反事实推理。