For a given causal question, it is important to efficiently decide which causal inference method to use for a given dataset. This is challenging because causal methods typically rely on complex and difficult-to-verify assumptions, and cross-validation is not applicable since ground truth causal quantities are unobserved.In this work, we propose CAusal Method Predictor (CAMP), a framework for predicting the best method for a given dataset. To this end, we generate datasets from a diverse set of synthetic causal models, score the candidate methods, and train a model to directly predict the highest-scoring method for that dataset. Next, by formulating a self-supervised pre-training objective centered on dataset assumptions relevant for causal inference, we significantly reduce the need for costly labeled data and enhance training efficiency. Our strategy learns to map implicit dataset properties to the best method in a data-driven manner. In our experiments, we focus on method prediction for causal discovery. CAMP outperforms selecting any individual candidate method and demonstrates promising generalization to unseen semi-synthetic and real-world benchmarks.
翻译:对于特定的因果问题,如何针对给定数据集高效选择适用的因果推断方法至关重要。这一任务极具挑战性,因为因果方法通常依赖于复杂且难以验证的假设,而交叉验证亦不适用——真实因果量无法被观测。
本文提出因果方法预测器(CAMP)框架,用于预测给定数据集的最优方法。为此,我们基于多样化合成因果模型生成数据集,对候选方法进行评分,并训练模型直接预测该数据集的最高评分方法。通过设计面向因果推断相关数据集假设的自监督预训练目标,我们显著降低了对昂贵标注数据的需求,并提升了训练效率。该策略以数据驱动方式学习将隐式数据集属性映射至最优方法。实验中,我们聚焦于因果发现的预测方法。CAMP超越所有单一候选方法的选择,并展现出在未见过的半合成及真实世界基准测试中的优异泛化能力。