Identifying latent variables and causal structures from observational data is essential to many real-world applications involving biological data, medical data, and unstructured data such as images and languages. However, this task can be highly challenging, especially when observed variables are generated by causally related latent variables and the relationships are nonlinear. In this work, we investigate the identification problem for nonlinear latent hierarchical causal models in which observed variables are generated by a set of causally related latent variables, and some latent variables may not have observed children. We show that the identifiability of causal structures and latent variables (up to invertible transformations) can be achieved under mild assumptions: on causal structures, we allow for multiple paths between any pair of variables in the graph, which relaxes latent tree assumptions in prior work; on structural functions, we permit general nonlinearity and multi-dimensional continuous variables, alleviating existing work's parametric assumptions. Specifically, we first develop an identification criterion in the form of novel identifiability guarantees for an elementary latent variable model. Leveraging this criterion, we show that both causal structures and latent variables of the hierarchical model can be identified asymptotically by explicitly constructing an estimation procedure. To the best of our knowledge, our work is the first to establish identifiability guarantees for both causal structures and latent variables in nonlinear latent hierarchical models.
翻译:从观测数据中识别潜在变量和因果结构对许多实际应用至关重要,涉及生物数据、医学数据以及图像和语言等非结构化数据。然而,这一任务极具挑战性,尤其是当观测变量由具有因果关系的潜在变量生成且关系为非线性时。本研究探讨了非线性潜变量层次因果模型的识别问题:其中观测变量由一组具有因果关系的潜在变量生成,且部分潜在变量可能没有可观测的子节点。我们证明,在温和假设下可实现因果结构和潜在变量(至可逆变换)的可辨识性:在因果结构方面,我们允许图中任意变量对之间存在多条路径,这放宽了先前工作中对潜在树形结构的假设;在结构函数方面,我们允许一般非线性与多维连续变量,从而弱化了现有工作的参数化假设。具体而言,首先针对基本潜变量模型提出一种以新型可辨识性保证为形式的识别准则。利用该准则,我们通过显式构建估计过程证明层次模型的因果结构和潜在变量均可渐近识别。据我们所知,本研究首次建立了非线性潜变量层次模型中因果结构与潜在变量的可辨识性保证。