Interpretability and transparency are essential for incorporating causal effect models from observational data into policy decision-making. They can provide trust for the model in the absence of ground truth labels to evaluate the accuracy of such models. To date, attempts at transparent causal effect estimation consist of applying post hoc explanation methods to black-box models, which are not interpretable. Here, we present BICauseTree: an interpretable balancing method that identifies clusters where natural experiments occur locally. Our approach builds on decision trees with a customized objective function to improve balancing and reduce treatment allocation bias. Consequently, it can additionally detect subgroups presenting positivity violations, exclude them, and provide a covariate-based definition of the target population we can infer from and generalize to. We evaluate the method's performance using synthetic and realistic datasets, explore its bias-interpretability tradeoff, and show that it is comparable with existing approaches.
翻译:可解释性与透明度对于将观测数据中的因果效应模型纳入政策决策至关重要,它们能在缺乏真实标签评估模型准确性时提供对模型的信任。然而目前实现透明因果效应估计的尝试主要局限于对黑箱模型采用事后解释方法,这类方法本身不具备可解释性。本文提出BICauseTree:一种可解释的平衡方法,能够识别局部自然实验发生的聚类簇。该方法基于决策树构建,通过定制化目标函数提升平衡性并减少处理分配偏差,进而可额外检测违背正向性假设的子群并进行排除,同时提供可推断与泛化的目标人群协变量定义。我们通过合成数据集与真实数据集评估方法性能,探索其偏差-可解释性权衡,并证明该方法达到与现有方法可比的水平。