Although understanding and characterizing causal effects have become essential in observational studies, it is challenging when the confounders are high-dimensional. In this article, we develop a general framework $\textit{CausalEGM}$ for estimating causal effects by encoding generative modeling, which can be applied in both binary and continuous treatment settings. Under the potential outcome framework with unconfoundedness, we establish a bidirectional transformation between the high-dimensional confounders space and a low-dimensional latent space where the density is known (e.g., multivariate normal distribution). Through this, CausalEGM simultaneously decouples the dependencies of confounders on both treatment and outcome and maps the confounders to the low-dimensional latent space. By conditioning on the low-dimensional latent features, CausalEGM can estimate the causal effect for each individual or the average causal effect within a population. Our theoretical analysis shows that the excess risk for CausalEGM can be bounded through empirical process theory. Under an assumption on encoder-decoder networks, the consistency of the estimate can be guaranteed. In a series of experiments, CausalEGM demonstrates superior performance over existing methods for both binary and continuous treatments. Specifically, we find CausalEGM to be substantially more powerful than competing methods in the presence of large sample sizes and high dimensional confounders. The software of CausalEGM is freely available at https://github.com/SUwonglab/CausalEGM.


翻译:尽管理解和刻画因果效应在观察性研究中已变得至关重要,但当混杂变量为高维时,这一任务极具挑战性。本文提出了一种通用框架 $\textit{CausalEGM}$,通过编码生成建模估计因果效应,可同时适用于二值处理和连续处理场景。在无混淆假设下的潜结果框架中,我们在高维混杂变量空间与密度已知(如多元正态分布)的低维潜空间之间建立了双向变换。通过这一变换,CausalEGM同时解耦了混杂变量对处理变量和结果变量的依赖关系,并将混杂变量映射至低维潜空间。通过以低维潜特征为条件,CausalEGM可估计每个个体的因果效应或群体平均因果效应。理论分析表明,基于经验过程理论,CausalEGM的超额风险可被界定量化。在编码器-解码器网络的假设下,估计的一致性得以保证。通过一系列实验,CausalEGM在二值处理和连续处理场景中均展现出优于现有方法的性能。具体而言,当样本量较大且混杂变量维度较高时,CausalEGM的效能显著强于竞争方法。CausalEGM的软件可在 https://github.com/SUwonglab/CausalEGM 自由获取。

0
下载
关闭预览

相关内容

因果推断,Causal Inference:The Mixtape
专知会员服务
110+阅读 · 2021年8月27日
近期必读的六篇 NeurIPS 2020【因果推理】相关论文和代码
专知会员服务
72+阅读 · 2020年10月31日
因果图,Causal Graphs,52页ppt
专知会员服务
254+阅读 · 2020年4月19日
因果效应估计组合拳:Reweighting和Representation
PaperWeekly
0+阅读 · 2022年9月2日
Hierarchically Structured Meta-learning
CreateAMind
27+阅读 · 2019年5月22日
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
深度自进化聚类:Deep Self-Evolution Clustering
我爱读PAMI
15+阅读 · 2019年4月13日
A Technical Overview of AI & ML in 2018 & Trends for 2019
待字闺中
18+阅读 · 2018年12月24日
【论文】变分推断(Variational inference)的总结
机器学习研究会
39+阅读 · 2017年11月16日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
26+阅读 · 2011年12月31日
国家自然科学基金
1+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
Arxiv
29+阅读 · 2023年2月10日
Arxiv
12+阅读 · 2022年11月21日
Arxiv
14+阅读 · 2022年10月15日
Arxiv
15+阅读 · 2022年1月24日
Arxiv
12+阅读 · 2021年6月29日
Arxiv
15+阅读 · 2020年12月17日
Arxiv
113+阅读 · 2020年2月5日
VIP会员
最新内容
博士论文 | 面向大模型推理的内存高效算法
专知会员服务
0+阅读 · 今天15:20
美空军新型反无人机部队初探
专知会员服务
4+阅读 · 今天5:45
《防空交战流程的概率建模研究》
专知会员服务
6+阅读 · 今天5:04
ICML 2026 教程 | 数值优化理论还重要吗?
专知会员服务
4+阅读 · 7月26日
ICM 2026 | 陶哲轩:人工智能时代的数学
专知会员服务
8+阅读 · 7月26日
《反无人机交战场景下的战斗归零研究》
专知会员服务
7+阅读 · 7月26日
博士论文 | 用代码结构感知方法推进代码大模型
相关论文
Arxiv
29+阅读 · 2023年2月10日
Arxiv
12+阅读 · 2022年11月21日
Arxiv
14+阅读 · 2022年10月15日
Arxiv
15+阅读 · 2022年1月24日
Arxiv
12+阅读 · 2021年6月29日
Arxiv
15+阅读 · 2020年12月17日
Arxiv
113+阅读 · 2020年2月5日
相关基金
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
26+阅读 · 2011年12月31日
国家自然科学基金
1+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员