Sample, estimate, aggregate: A recipe for causal discovery foundation models

Causal discovery, the task of inferring causal structure from data, has the potential to uncover mechanistic insights from biological experiments, especially those involving perturbations. However, causal discovery algorithms over larger sets of variables tend to be brittle against misspecification or when data are limited. For example, single-cell transcriptomics measures thousands of genes, but the nature of their relationships is not known, and there may be as few as tens of cells per intervention setting. To mitigate these challenges, we propose a foundation model-inspired approach: a supervised model trained on large-scale, synthetic data to predict causal graphs from summary statistics -- like the outputs of classical causal discovery algorithms run over subsets of variables and other statistical hints like inverse covariance. Our approach is enabled by the observation that typical errors in the outputs of a discovery algorithm remain comparable across datasets. Theoretically, we show that the model architecture is well-specified, in the sense that it can recover a causal graph consistent with graphs over subsets. Empirically, we train the model to be robust to misspecification and distribution shift using diverse datasets. Experiments on biological and synthetic data confirm that this model generalizes well beyond its training set, runs on graphs with hundreds of variables in seconds, and can be easily adapted to different underlying data assumptions.

翻译：因果发现是从数据中推断因果结构的任务，具有从生物学实验（特别是涉及干预的实验）中揭示机制性见解的潜力。然而，针对较大变量集的因果发现算法往往在模型设定错误或数据有限时表现脆弱。例如，单细胞转录组学可测量数千个基因，但其相互关系本质未知，且每个干预条件下的细胞数量可能仅有数十个。为缓解这些挑战，我们提出一种受基础模型启发的监督学习方法：该模型在大规模合成数据上训练，能够根据汇总统计量（如在变量子集上运行经典因果发现算法得到的输出，以及逆协方差等其他统计线索）预测因果图。该方法基于以下观察得以实现：发现算法输出中的典型误差在不同数据集间保持可比性。理论上，我们证明该模型架构具有良好设定性，即能够恢复与子图一致的因果图。实证方面，我们通过多样化数据集训练模型，使其对设定错误和分布偏移具有鲁棒性。在生物与合成数据上的实验表明，该模型能良好泛化至训练集之外，可在数秒内处理包含数百个变量的图结构，并能轻松适配不同的底层数据假设。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

《生成式模型: 变分自编码器与扩散模型》，75页ppt，Google DeepMind科学家Ruiqi Gao

专知会员服务

66+阅读 · 2023年6月10日

《用于无线通信和传感的智能反射面 (IRS)》（ICC 2022）新加坡国立大学2022最新53页slides

专知会员服务

25+阅读 · 2022年11月16日

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

分布外泛化(Out-Of-Distribution Generalization) 综述论文，22页pdf240篇文献

专知会员服务

64+阅读 · 2021年9月2日