Recent years have experienced increasing utilization of complex machine learning models across multiple sources of data to inform more generalizable decision-making. However, distribution shifts across data sources and privacy concerns related to sharing individual-level data, coupled with a lack of uncertainty quantification from machine learning predictions, make it challenging to achieve valid inferences in multi-source environments. In this paper, we consider the problem of obtaining distribution-free prediction intervals for a target population, leveraging multiple potentially biased data sources. We derive the efficient influence functions for the quantiles of unobserved outcomes in the target and source populations, and show that one can incorporate machine learning prediction algorithms in the estimation of nuisance functions while still achieving parametric rates of convergence to nominal coverage probabilities. Moreover, when conditional outcome invariance is violated, we propose a data-adaptive strategy to upweight informative data sources for efficiency gain and downweight non-informative data sources for bias reduction. We highlight the robustness and efficiency of our proposals for a variety of conformal scores and data-generating mechanisms via extensive synthetic experiments. Hospital length of stay prediction intervals for pediatric patients undergoing a high-risk cardiac surgical procedure between 2016-2022 in the U.S. illustrate the utility of our methodology.
翻译:近年来,利用跨多个数据源的复杂机器学习模型以支持更具泛化性的决策制定愈发普遍。然而,数据源间的分布偏移、共享个体层面数据的隐私顾虑,以及机器学习预测缺乏不确定性量化等问题,使得在多源环境中实现有效推断面临挑战。本文研究从多个潜在有偏数据源出发,为目标总体构建无分布假设预测区间的问题。我们推导出目标总体与源总体中未观测结果分位数的高效影响函数,并证明在估计干扰函数时可融入机器学习预测算法,同时仍能达到名义覆盖概率的参数收敛速率。此外,当条件结果不变性被违反时,我们提出一种数据自适应策略:对信息量大的数据源进行上加权以提高效率,对非信息性数据源进行下加权以减少偏倚。通过大量合成实验,我们展示了所提方法对多种共形分数与数据生成机制的稳健性与高效性。针对2016-2022年美国接受高风险心脏手术的儿科患者住院时长预测区间的实证分析,进一步证明了该方法的实用价值。