We address the problem of integrating data from multiple, possibly biased, observational and interventional studies, to eventually compute counterfactuals in structural causal models. We start from the case of a single observational dataset affected by a selection bias. We show that the likelihood of the available data has no local maxima. This enables us to use the causal expectation-maximisation scheme to approximate the bounds for partially identifiable counterfactual queries, which are the focus of this paper. We then show how the same approach can address the general case of multiple datasets, no matter whether interventional or observational, biased or unbiased, by remapping it into the former one via graphical transformations. Systematic numerical experiments and a case study on palliative care show the effectiveness of our approach, while hinting at the benefits of fusing heterogeneous data sources to get informative outcomes in case of partial identifiability.
翻译:我们解决了整合来自多个可能带有偏倚的观测和干预研究数据的问题,以最终在结构因果模型中计算反事实。我们从受选择偏倚影响的单个观测数据集的情况入手。我们表明,可用数据的似然函数不存在局部极大值。这使我们能够利用因果期望最大化方案来近似部分可识别反事实查询的界,这是本文的重点。然后,我们展示了如何通过图形变换将相同的方法重新映射到前一种情况,从而处理多个数据集的一般情况,无论这些数据集是干预性还是观测性、偏倚还是无偏倚。系统的数值实验和一项关于姑息治疗的案例研究显示了我们的方法的有效性,同时暗示了融合异质数据源在部分可识别情况下获取信息丰富结果的好处。