A primary challenge facing modern scientific research is the limited availability of gold-standard data which can be both costly and labor-intensive to obtain. With the rapid development of machine learning (ML), scientists have relied on ML algorithms to predict these gold-standard outcomes with easily obtained covariates. However, these predicted outcomes are often used directly in subsequent statistical analyses, ignoring imprecision and heterogeneity introduced by the prediction procedure. This will likely result in false positive findings and invalid scientific conclusions. In this work, we introduce an assumption-lean and data-adaptive Post-Prediction Inference (POP-Inf) procedure that allows valid and powerful inference based on ML-predicted outcomes. Its "assumption-lean" property guarantees reliable statistical inference without assumptions on the ML-prediction, for a wide range of statistical quantities. Its "data-adaptive'" feature guarantees an efficiency gain over existing post-prediction inference methods, regardless of the accuracy of ML-prediction. We demonstrate the superiority and applicability of our method through simulations and large-scale genomic data.
翻译:现代科学研究面临的主要挑战之一是金标准数据的有限可用性,这类数据获取成本高昂且劳动密集。随着机器学习的快速发展,科学家们依赖机器学习算法利用易获取的协变量来预测这些金标准结果。然而,这些预测结果往往被直接用于后续统计分析,忽略了预测过程引入的不精确性和异质性,这很可能导致假阳性发现和无效的科学结论。在本研究中,我们提出了一种假设精简且数据自适应的事后预测推断程序,该程序能够基于机器学习预测结果进行有效且有力的推断。其“假设精简”特性保证了在不对机器学习预测施加假设的情况下,对广泛统计量进行可靠的统计推断。其“数据自适应”特性则确保无论机器学习预测的准确性如何,都能相较于现有事后预测推断方法实现效率提升。我们通过模拟实验和大规模基因组数据验证了该方法的优越性和适用性。