Causal discovery in the presence of unobserved common causes from observational data only is a crucial but challenging problem. We categorize all possible causal relationships between two random variables into the following four categories and aim to identify one from observed data: two cases in which either of the direct causality exists, a case that variables are independent, and a case that variables are confounded by latent confounders. Although existing methods have been proposed to tackle this problem, they require unobserved variables to satisfy assumptions on the form of their equation models. In our previous study (Kobayashi et al., 2022), the first causal discovery method without such assumptions is proposed for discrete data and named CLOUD. Using Normalized Maximum Likelihood (NML) Code, CLOUD selects a model that yields the minimum codelength of the observed data from a set of model candidates. This paper extends CLOUD to apply for various data types across discrete, mixed, and continuous. We not only performed theoretical analysis to show the consistency of CLOUD in terms of the model selection, but also demonstrated that CLOUD is more effective than existing methods in inferring causal relationships by extensive experiments on both synthetic and real-world data.
翻译:仅从观测数据中检测未观测共同因子存在下的因果关系是极具挑战性的关键问题。我们将两个随机变量间所有可能的因果关系归为四类:存在单向直接因果关系的两种情况、变量独立的情况、以及变量受潜在混杂因素混淆的情况,并致力于从观测数据中识别具体类型。现有方法虽能解决此问题,但要求未观测变量满足方程模型形式的特定假设。在前期研究(Kobayashi等人,2022)中,我们提出了首个无需此类假设的离散数据因果发现方法CLOUD。该方法采用归一化最大似然(NML)编码,从候选模型集合中选择使观测数据编码长度最短的模型。本文扩展了CLOUD方法,使其适用于离散、混合及连续等多种数据类型。我们不仅通过理论分析证明了CLOUD在模型选择上的一致性,还基于合成数据与真实数据的大量实验表明,CLOUD在推断因果关系方面优于现有方法。