For a given causal question, it is important to efficiently decide which causal inference method to use for a given dataset. This is challenging because causal methods typically rely on complex and difficult-to-verify assumptions, and cross-validation is not applicable since ground truth causal quantities are unobserved. In this work, we propose CAusal Method Predictor (CAMP), a framework for predicting the best method for a given dataset. To this end, we generate datasets from a diverse set of synthetic causal models, score the candidate methods, and train a model to directly predict the highest-scoring method for that dataset. Next, by formulating a self-supervised pre-training objective centered on dataset assumptions relevant for causal inference, we significantly reduce the need for costly labeled data and enhance training efficiency. Our strategy learns to map implicit dataset properties to the best method in a data-driven manner. In our experiments, we focus on method prediction for causal discovery. CAMP outperforms selecting any individual candidate method and demonstrates promising generalization to unseen semi-synthetic and real-world benchmarks.
翻译:对于给定的因果问题,高效地决定针对特定数据集应采用哪种因果推断方法至关重要。这一过程颇具挑战性,因为因果方法通常依赖于复杂且难以验证的假设,同时由于真实因果量未被观测,交叉验证方法无法适用。本研究提出因果方法预测器(CAMP)框架,用于预测特定数据集的最优方法。为此,我们通过多样化的合成因果模型生成数据集,对候选方法进行评分,并训练模型直接预测得分最高的方法。接着,通过构建以因果推断相关数据集假设为核心的自我监督预训练目标,我们显著降低了对昂贵标注数据的依赖,同时提升了训练效率。该策略以数据驱动方式学习将隐式数据集属性映射至最优方法。实验聚焦于因果发现的方法预测任务,CAMP不仅优于所有单个候选方法,还展现出对未见半合成数据集及真实世界基准测试的有效泛化能力。