In recent years, there has been a growing interest in using machine learning techniques for the estimation of treatment effects. Most of the best-performing methods rely on representation learning strategies that encourage shared behavior among potential outcomes to increase the precision of treatment effect estimates. In this paper we discuss and classify these models in terms of their algorithmic inductive biases and present a new model, NN-CGC, that considers additional information from the causal graph. NN-CGC tackles bias resulting from spurious variable interactions by implementing novel constraints on models, and it can be integrated with other representation learning methods. We test the effectiveness of our method using three different base models on common benchmarks. Our results indicate that our model constraints lead to significant improvements, achieving new state-of-the-art results in treatment effects estimation. We also show that our method is robust to imperfect causal graphs and that using partial causal information is preferable to ignoring it.
翻译:近年来,利用机器学习技术进行处理效应估计的研究日益受到关注。多数性能最优的方法依赖于表征学习策略,通过促进潜在结果间的共享行为来提高处理效应估计的精度。本文对这些模型进行系统性分类,并从算法归纳偏置角度展开讨论,提出一种考虑因果图额外信息的新模型NN-CGC。该模型通过引入新颖的约束机制来消除虚假变量交互导致的偏差,并可与其他表征学习方法集成。我们在通用基准上使用三种不同基础模型检验了方法的有效性。结果表明,模型约束带来了显著改进,在处理效应估计中取得了领先水平的新成果。我们还证明该方法对不完美的因果图具有鲁棒性,且利用部分因果信息优于完全忽略因果信息。