An acyclic causal structure can be described using a directed acyclic graph (DAG) with arrows indicating causation. The task of learning this structure from data is known as "causal discovery." Diverse populations or changing environments can sometimes give rise to heterogeneous data. This heterogeneity can be thought of as a mixture model with multiple "sources," each exerting their own distinct signature on the observed variables. From this perspective, the source is a latent common cause for every observed variable. While some methods for causal discovery are able to work around unobserved confounding in special cases, the only known ways to deal with a global confounder (such as a latent class) involve parametric assumptions. Focusing on discrete observables, we demonstrate that globally confounded causal structures can still be identifiable without parametric assumptions, so long as the number of latent classes remains small relative to the size and sparsity of the underlying DAG.
翻译:有向无环图(DAG)可通过箭头表示因果关系,从而描述非循环的因果结构。从数据中学习这种结构的任务被称为“因果发现”。不同的群体或变化的环境有时会产生异质性数据。这种异质性可被视为具有多个“源”的混合模型,每个源都对观测变量施加其独特的影响。从这个角度看,源是每个观测变量的一个潜在共同原因。尽管某些因果发现方法能够在特殊情况下处理未观测的混杂因素,但已知处理全局混杂因素(如潜在类别)的唯一方法涉及参数化假设。针对离散观测变量,我们证明只要潜在类别的数量相对于底层DAG的规模和稀疏性保持较小,即使没有参数化假设,全局混杂的因果结构仍然可以是可识别的。