This paper demonstrates how to discover the whole causal graph from the second derivative of the log-likelihood in observational non-linear additive Gaussian noise models. Leveraging scalable machine learning approaches to approximate the score function $\nabla \log p(\mathbf{X})$, we extend the work of Rolland et al. (2022) that only recovers the topological order from the score and requires an expensive pruning step removing spurious edges among those admitted by the ordering. Our analysis leads to DAS (acronym for Discovery At Scale), a practical algorithm that reduces the complexity of the pruning by a factor proportional to the graph size. In practice, DAS achieves competitive accuracy with current state-of-the-art while being over an order of magnitude faster. Overall, our approach enables principled and scalable causal discovery, significantly lowering the compute bar.
翻译:本文展示了如何从观测数据的非线性加性高斯噪声模型中,通过对数似然的二阶导数发现整个因果图。借助可扩展的机器学习方法来近似分数函数 $\nabla \log p(\mathbf{X})$,我们扩展了 Rolland 等人 (2022) 的工作,该工作仅从分数中恢复拓扑序,并需要昂贵的剪枝步骤来移除排序所允许边中的虚假边。我们的分析提出了 DAS(即“可扩展发现”的缩写),一种实际算法,将剪枝的复杂度降低了与图大小成比例的因子。在实践中,DAS 在与当前最先进方法竞争精度的同时,速度提升了一个数量级以上。总体而言,我们的方法实现了原则性且可扩展的因果发现,显著降低了计算门槛。