While reliable data-driven decision-making hinges on high-quality labeled data, the acquisition of quality labels often involves laborious human annotations or slow and expensive scientific measurements. Machine learning is becoming an appealing alternative as sophisticated predictive techniques are being used to quickly and cheaply produce large amounts of predicted labels; e.g., predicted protein structures are used to supplement experimentally derived structures, predictions of socioeconomic indicators from satellite imagery are used to supplement accurate survey data, and so on. Since predictions are imperfect and potentially biased, this practice brings into question the validity of downstream inferences. We introduce cross-prediction: a method for valid inference powered by machine learning. With a small labeled dataset and a large unlabeled dataset, cross-prediction imputes the missing labels via machine learning and applies a form of debiasing to remedy the prediction inaccuracies. The resulting inferences achieve the desired error probability and are more powerful than those that only leverage the labeled data. Closely related is the recent proposal of prediction-powered inference, which assumes that a good pre-trained model is already available. We show that cross-prediction is consistently more powerful than an adaptation of prediction-powered inference in which a fraction of the labeled data is split off and used to train the model. Finally, we observe that cross-prediction gives more stable conclusions than its competitors; its confidence intervals typically have significantly lower variability.
翻译:尽管可靠的数据驱动决策依赖于高质量标注数据,但获取优质标注往往需要繁重的人工标注或缓慢且昂贵的科学测量。机器学习正成为一种颇具吸引力的替代方案——先进预测技术被用于快速廉价地生成大量预测标签;例如,利用预测的蛋白质结构补充实验解析结构,通过卫星图像预测社会经济指标补充精确调查数据等。由于预测结果存在不完善性和潜在偏差,这种做法使得下游推断的有效性存疑。我们提出交叉预测:一种借助机器学习实现有效推断的方法。通过少量标注数据集和大量未标注数据,交叉预测利用机器学习填补缺失标签,并采用去偏技术纠正预测误差。由此产生的推断能达到期望的错误概率,且比仅利用标注数据的方法更具统计效力。与此密切相关的是近期提出的预测驱动推断方法,其假设已存在优质的预训练模型。我们证明,相较于将部分标注数据拆分用于训练模型的预测驱动推断改编版本,交叉预测始终具有更高的统计效力。最后,我们观察到交叉预测能比竞争方法得出更稳定的结论,其置信区间的变异性通常显著更低。