基于文本嵌入与深度学习的因果估计解混淆方法 (Reading Between the Lines: Deconfounding Causal Estimates using Text Embeddings and Deep Learning) - 专知论文

会员服务 ·

0

嵌入 · 结构 · 文本嵌入 · 结构化 · 高维 ·

Reading Between the Lines: Deconfounding Causal Estimates using Text Embeddings and Deep Learning

翻译：基于文本嵌入与深度学习的因果估计解混淆方法

Ahmed Dawoud,Osama El-Shamy

Estimating causal treatment effects in observational settings is frequently compromised by selection bias arising from unobserved confounders. While traditional econometric methods struggle when these confounders are orthogonal to structured covariates, high-dimensional unstructured text often contains rich proxies for these latent variables. This study proposes a Neural Network-Enhanced Double Machine Learning (DML) framework designed to leverage text embeddings for causal identification. Using a rigorous synthetic benchmark, we demonstrate that unstructured text embeddings capture critical confounding information that is absent from structured tabular data. However, we show that standard tree-based DML estimators retain substantial bias (+24%) due to their inability to model the continuous topology of embedding manifolds. In contrast, our deep learning approach reduces bias to -0.86% with optimized architectures, effectively recovering the ground-truth causal parameter. These findings suggest that deep learning architectures are essential for satisfying the unconfoundedness assumption when conditioning on high-dimensional natural language data

翻译：在观测性研究中估计因果处理效应常因未观测混杂因素导致的选择偏倚而失真。当这些混杂因素与结构化协变量正交时，传统计量经济学方法往往失效，而高维非结构化文本数据通常蕴含这些潜变量的丰富代理信息。本研究提出一种神经网络增强的双重机器学习框架，旨在利用文本嵌入实现因果识别。通过严格的合成基准测试，我们证明非结构化文本嵌入能捕获结构化表格数据中缺失的关键混杂信息。然而，我们发现标准基于树的DML估计量因无法建模嵌入流形的连续拓扑结构而存在显著偏倚（+24%）。相比之下，我们的深度学习方法通过优化架构将偏倚降低至-0.86%，有效恢复了真实因果参数。这些发现表明，当以高维自然语言数据为条件时，深度学习架构对于满足无混杂性假设具有不可或缺的作用。

0

相关内容

西安交大最新《深度学习因果模型》综述论文，35页pdf涵盖292篇文献阐述三种数据范式因果模型

西安交大最新《深度学习因果模型》综述论文，35页pdf涵盖292篇文献阐述三种数据范式因果模型

专知会员服务

63+阅读 · 2023年11月5日

【CMU博士论文】非参数因果推断，241页pdf

【CMU博士论文】非参数因果推断，241页pdf

专知会员服务

35+阅读 · 2023年6月20日

【AISTATS2023】基于上下文和混杂因素的因果效应估计，77页ppt

【AISTATS2023】基于上下文和混杂因素的因果效应估计，77页ppt

专知会员服务

30+阅读 · 2023年4月29日

清华等最新《因果强化学习》综述，29页pdf详述因果强化学习方法与评价

清华等最新《因果强化学习》综述，29页pdf详述因果强化学习方法与评价

专知会员服务

102+阅读 · 2023年2月13日

【剑桥大学博士论文】监督学习、模仿和强化学习中泛化和自适应的因果表示学习，202页pdf

【剑桥大学博士论文】监督学习、模仿和强化学习中泛化和自适应的因果表示学习，202页pdf

专知会员服务

54+阅读 · 2023年2月3日

深度学习和因果如何结合？北交最新《深度因果模型》综述论文，31页pdf涵盖216页pdf详述41个深度因果模型

深度学习和因果如何结合？北交最新《深度因果模型》综述论文，31页pdf涵盖216页pdf详述41个深度因果模型

专知会员服务

128+阅读 · 2022年9月21日

西安交大最新《深度学习因果发现》综述论文，26页pdf涵盖211篇文献阐述三种深度因果范式

西安交大最新《深度学习因果发现》综述论文，26页pdf涵盖211篇文献阐述三种深度因果范式

专知会员服务

117+阅读 · 2022年9月15日

什么是因果深度学习？DeepMind最新ICML2022《因果性与深度学习:协同、挑战和未来》教程，183页ppt详述因果DL

什么是因果深度学习？DeepMind最新ICML2022《因果性与深度学习:协同、挑战和未来》教程，183页ppt详述因果DL

专知会员服务

200+阅读 · 2022年7月20日

最新「因果推断Causal Inference」综述论文38页pdf，Buffalo、Georgia、阿里巴巴、Virginia

专知会员服务

183+阅读 · 2020年2月11日

【独立研究者I-Sheng Yang论文】因果机器学习损失函数（A Loss-Function for Causal Machine-Learning）

【独立研究者I-Sheng Yang论文】因果机器学习损失函数（A Loss-Function for Causal Machine-Learning）

专知会员服务

20+阅读 · 2020年1月7日

「因果推理」概述论文，13页pdf

「因果推理」概述论文，13页pdf

专知

16+阅读 · 2021年3月20日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知

11+阅读 · 2020年8月28日

基于深度元学习的因果推断新方法

基于深度元学习的因果推断新方法

图与推荐

12+阅读 · 2020年7月21日

【ICML2020-Tutorial】因果强化学习-CRL，147页ppt，哥伦比亚大学-Elias Bareinboim

【ICML2020-Tutorial】因果强化学习-CRL，147页ppt，哥伦比亚大学-Elias Bareinboim

专知

13+阅读 · 2020年7月16日

AAAI2020最新「因果推理表示学习」122页ppt，Georgia、Buffalo、阿里巴巴与Virginia

AAAI2020最新「因果推理表示学习」122页ppt，Georgia、Buffalo、阿里巴巴与Virginia

专知

16+阅读 · 2020年2月12日

最新「因果推断Causal Inference」综述论文38页pdf，阿里巴巴、Buffalo、Georgia、Virginia

最新「因果推断Causal Inference」综述论文38页pdf，阿里巴巴、Buffalo、Georgia、Virginia

专知

68+阅读 · 2020年2月11日

用深度学习揭示数据的因果关系

用深度学习揭示数据的因果关系

专知

28+阅读 · 2019年5月18日

因果推理学习算法资源大列表

因果推理学习算法资源大列表

专知

27+阅读 · 2019年3月3日

让智能体主动交互，DeepMind提出用元强化学习实现因果推理

让智能体主动交互，DeepMind提出用元强化学习实现因果推理

机器之心

17+阅读 · 2019年2月11日

【论文】深度学习的数学解释

【论文】深度学习的数学解释

机器学习研究会

10+阅读 · 2017年12月15日

图文混合跨媒体知识单元的模糊分类方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

测量误差数据下部分线性模型有约束统计推断理论

国家自然科学基金

2+阅读 · 2015年12月31日

复杂环境下机器学习的理论研究

国家自然科学基金

21+阅读 · 2015年12月31日

多标记文本数据流分类方法研究

国家自然科学基金

3+阅读 · 2015年12月31日

基于生态演替的文本大数据特征学习研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于相依数据的梯度学习理论研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于深度学习的机器译文质量估计方法研究

国家自然科学基金

3+阅读 · 2014年12月31日

含有隐变量的因果结构学习与统计因果推断

国家自然科学基金

21+阅读 · 2013年12月31日

因果推断的统计方法

国家自然科学基金

26+阅读 · 2011年12月31日

因果推断及不完全数据的统计分析

国家自然科学基金

23+阅读 · 2008年12月31日

Information-Theoretic Causal Bounds under Unmeasured Confounding

Arxiv

0+阅读 · 2月3日

Disentangling spatial interference and spatial confounding biases in causal inference

Arxiv

0+阅读 · 2月2日

Data-Driven Information-Theoretic Causal Bounds under Unmeasured Confounding

Arxiv

0+阅读 · 1月23日

Many Experiments, Few Repetitions, Unpaired Data, and Sparse Effects: Is Causal Inference Possible?

Arxiv

0+阅读 · 1月21日

Lost in Aggregation: The Causal Interpretation of the IV Estimand

Arxiv

0+阅读 · 1月17日

Reevaluating Causal Estimation Methods with Data from a Product Release

Arxiv

0+阅读 · 1月17日

Coupling Generative Modeling and an Autoencoder with the Causal Bridge

Arxiv

0+阅读 · 1月14日

Estimating Causal Effects in Gaussian Linear SCMs with Finite Data

Arxiv

0+阅读 · 1月8日

Causal Discovery with Mixed Latent Confounding via Precision Decomposition

Arxiv

0+阅读 · 2025年12月31日

Estimation and Inference for Causal Explainability

Arxiv

0+阅读 · 2025年12月30日

VIP会员

文章信息

相关主题

相关VIP内容

西安交大最新《深度学习因果模型》综述论文，35页pdf涵盖292篇文献阐述三种数据范式因果模型

西安交大最新《深度学习因果模型》综述论文，35页pdf涵盖292篇文献阐述三种数据范式因果模型

专知会员服务

63+阅读 · 2023年11月5日

【CMU博士论文】非参数因果推断，241页pdf

【CMU博士论文】非参数因果推断，241页pdf

专知会员服务

35+阅读 · 2023年6月20日

【AISTATS2023】基于上下文和混杂因素的因果效应估计，77页ppt

【AISTATS2023】基于上下文和混杂因素的因果效应估计，77页ppt

专知会员服务

30+阅读 · 2023年4月29日

清华等最新《因果强化学习》综述，29页pdf详述因果强化学习方法与评价

清华等最新《因果强化学习》综述，29页pdf详述因果强化学习方法与评价

专知会员服务

102+阅读 · 2023年2月13日

【剑桥大学博士论文】监督学习、模仿和强化学习中泛化和自适应的因果表示学习，202页pdf

【剑桥大学博士论文】监督学习、模仿和强化学习中泛化和自适应的因果表示学习，202页pdf

专知会员服务

54+阅读 · 2023年2月3日

深度学习和因果如何结合？北交最新《深度因果模型》综述论文，31页pdf涵盖216页pdf详述41个深度因果模型

深度学习和因果如何结合？北交最新《深度因果模型》综述论文，31页pdf涵盖216页pdf详述41个深度因果模型

专知会员服务

128+阅读 · 2022年9月21日

西安交大最新《深度学习因果发现》综述论文，26页pdf涵盖211篇文献阐述三种深度因果范式

西安交大最新《深度学习因果发现》综述论文，26页pdf涵盖211篇文献阐述三种深度因果范式

专知会员服务

117+阅读 · 2022年9月15日

什么是因果深度学习？DeepMind最新ICML2022《因果性与深度学习:协同、挑战和未来》教程，183页ppt详述因果DL

什么是因果深度学习？DeepMind最新ICML2022《因果性与深度学习:协同、挑战和未来》教程，183页ppt详述因果DL

专知会员服务

200+阅读 · 2022年7月20日

最新「因果推断Causal Inference」综述论文38页pdf，Buffalo、Georgia、阿里巴巴、Virginia

专知会员服务

183+阅读 · 2020年2月11日

【独立研究者I-Sheng Yang论文】因果机器学习损失函数（A Loss-Function for Causal Machine-Learning）

【独立研究者I-Sheng Yang论文】因果机器学习损失函数（A Loss-Function for Causal Machine-Learning）

专知会员服务

20+阅读 · 2020年1月7日

热门VIP内容

开通专知VIP会员享更多权益服务

《无人机与战争：被忽视的环境影响及无人机保护潜力》

俄罗斯规划未来无人机驱动军队

《整合杀伤链：一个用于边缘目标验证与战术推理的零样本框架》最新资料

《人工智能、武器与影响力：前沿模型在模拟核危机中展现复杂推理》2026最新46页报告

相关资讯

「因果推理」概述论文，13页pdf

「因果推理」概述论文，13页pdf

专知

16+阅读 · 2021年3月20日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知

11+阅读 · 2020年8月28日

基于深度元学习的因果推断新方法

基于深度元学习的因果推断新方法

图与推荐

12+阅读 · 2020年7月21日

【ICML2020-Tutorial】因果强化学习-CRL，147页ppt，哥伦比亚大学-Elias Bareinboim

【ICML2020-Tutorial】因果强化学习-CRL，147页ppt，哥伦比亚大学-Elias Bareinboim

专知

13+阅读 · 2020年7月16日

AAAI2020最新「因果推理表示学习」122页ppt，Georgia、Buffalo、阿里巴巴与Virginia

AAAI2020最新「因果推理表示学习」122页ppt，Georgia、Buffalo、阿里巴巴与Virginia

专知

16+阅读 · 2020年2月12日

最新「因果推断Causal Inference」综述论文38页pdf，阿里巴巴、Buffalo、Georgia、Virginia

最新「因果推断Causal Inference」综述论文38页pdf，阿里巴巴、Buffalo、Georgia、Virginia

专知

68+阅读 · 2020年2月11日

用深度学习揭示数据的因果关系

用深度学习揭示数据的因果关系

专知

28+阅读 · 2019年5月18日

因果推理学习算法资源大列表

因果推理学习算法资源大列表

专知

27+阅读 · 2019年3月3日

让智能体主动交互，DeepMind提出用元强化学习实现因果推理

让智能体主动交互，DeepMind提出用元强化学习实现因果推理

机器之心

17+阅读 · 2019年2月11日

【论文】深度学习的数学解释

【论文】深度学习的数学解释

机器学习研究会

10+阅读 · 2017年12月15日

相关论文

Information-Theoretic Causal Bounds under Unmeasured Confounding

Arxiv

0+阅读 · 2月3日

Disentangling spatial interference and spatial confounding biases in causal inference

Arxiv

0+阅读 · 2月2日

Data-Driven Information-Theoretic Causal Bounds under Unmeasured Confounding

Arxiv

0+阅读 · 1月23日

Many Experiments, Few Repetitions, Unpaired Data, and Sparse Effects: Is Causal Inference Possible?

Arxiv

0+阅读 · 1月21日

Lost in Aggregation: The Causal Interpretation of the IV Estimand

Arxiv

0+阅读 · 1月17日

Reevaluating Causal Estimation Methods with Data from a Product Release

Arxiv

0+阅读 · 1月17日

Coupling Generative Modeling and an Autoencoder with the Causal Bridge

Arxiv

0+阅读 · 1月14日

Estimating Causal Effects in Gaussian Linear SCMs with Finite Data

Arxiv

0+阅读 · 1月8日

Causal Discovery with Mixed Latent Confounding via Precision Decomposition

Arxiv

0+阅读 · 2025年12月31日

Estimation and Inference for Causal Explainability

Arxiv

0+阅读 · 2025年12月30日

相关基金

图文混合跨媒体知识单元的模糊分类方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

测量误差数据下部分线性模型有约束统计推断理论

国家自然科学基金

2+阅读 · 2015年12月31日

复杂环境下机器学习的理论研究

国家自然科学基金

21+阅读 · 2015年12月31日

多标记文本数据流分类方法研究

国家自然科学基金

3+阅读 · 2015年12月31日

基于生态演替的文本大数据特征学习研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于相依数据的梯度学习理论研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于深度学习的机器译文质量估计方法研究

国家自然科学基金

3+阅读 · 2014年12月31日

含有隐变量的因果结构学习与统计因果推断

国家自然科学基金

21+阅读 · 2013年12月31日

因果推断的统计方法

国家自然科学基金

26+阅读 · 2011年12月31日

因果推断及不完全数据的统计分析

国家自然科学基金

23+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员