Estimating the conditional average treatment effect (CATE) from observational data is relevant for many applications such as personalized medicine. Here, we focus on the widespread setting where the observational data come from multiple environments, such as different hospitals, physicians, or countries. Furthermore, we allow for violations of standard causal assumptions, namely, overlap within the environments and unconfoundedness. To this end, we move away from point identification and focus on partial identification. Specifically, we show that current assumptions from the literature on multiple environments allow us to interpret the environment as an instrumental variable (IV). This allows us to adapt bounds from the IV literature for partial identification of CATE by leveraging treatment assignment mechanisms across environments. Then, we propose different model-agnostic learners (so-called meta-learners) to estimate the bounds that can be used in combination with arbitrary machine learning models. We further demonstrate the effectiveness of our meta-learners across various experiments using both simulated and real-world data. Finally, we discuss the applicability of our meta-learners to partial identification in instrumental variable settings, such as randomized controlled trials with non-compliance.
翻译:从观测数据中估计条件平均处理效应(CATE)对于个性化医疗等众多应用具有重要意义。本文聚焦于观测数据来自多个环境(如不同医院、医生或国家)的普遍场景。此外,我们允许标准因果假设(即环境内的重叠性与无混淆性)被违反。为此,我们放弃点识别而专注于部分识别。具体而言,我们证明了当前多环境文献中的假设允许我们将环境解释为工具变量(IV)。通过利用跨环境的处理分配机制,这使我们能够调整工具变量文献中的边界以实现CATE的部分识别。随后,我们提出了多种与模型无关的学习器(即元学习器)来估计这些边界,这些学习器可与任意机器学习模型结合使用。我们进一步通过模拟数据和真实世界数据的多组实验验证了所提元学习器的有效性。最后,我们讨论了元学习器在工具变量场景(如存在不依从性的随机对照试验)中部分识别的适用性。