Attributing an observed outcome to its root cause is a central task in domains ranging from medical diagnosis to engineering fault diagnosis. Existing approaches either equate the root cause with a root node of the causal graph, as in causal-discovery-based root cause analysis, or target causes more broadly and thereby favour proximate ones, as with the probability of causation and posterior causal effects. We argue that this issue stems from the absence of a formal definition of a root cause, which has led to methods designed for other purposes being applied to root cause attribution by default. We address this by giving a formal, individual-level definition of a root cause within the potential outcomes framework, based on the notion of an individual cause and a counterfactual root condition motivated by mediation analysis. Building on this definition, we propose the probability of root cause (PRC), which quantifies how probable it is that a candidate variable set is the root cause of a given outcome, conditional on observed evidence. Under standard assumptions, we establish the identifiability of the PRC and derive an explicit identification formula. Two numerical examples illustrate the approach.
翻译:将观察到的结果归因于其根本原因是医学诊断到工程故障诊断等领域的核心任务。现有方法要么将根本原因等同于因果图中的根节点(如基于因果发现的根本原因分析),要么更宽泛地定位原因从而更倾向于近因(如因果关系概率和后验因果效应)。我们认为,这一问题源于缺乏对根本原因的正式定义,导致为其他目的设计的方法被默认用于根本原因归因。为此,我们在潜在结果框架内,基于个体原因概念和由中介分析启发的反事实根本条件,给出了根本原因的正式个体级定义。基于这一定义,我们提出了根本原因的概率(PRC),用以量化在给定观测证据条件下,候选变量集是特定结果根本原因的概率。在标准假设下,我们确立了PRC的可识别性,并推导出显式的识别公式。两个数值示例说明了该方法。