Unlabeled data are increasingly prevalent in contemporary economic studies, yet their effective use for improving prediction remains challenging because the outcomes are often costly or even infeasible to observe. Machine learning methods can help label these data and achieve high predictive accuracy, but they often lack interpretability. In this paper, we propose a Prediction-powered Unified Model Averaging (PUMA) framework to combine linear regression and machine learning methods, achieving a balance between interpretation and prediction. Unlike existing studies on prediction-powered inference, our approach is the first to jointly address uncertainty arising from model misspecification, power tuning parameter selection, and the choice of machine learning algorithms by using model averaging. Theoretically, under mild conditions, we establish the in-sample and out-of-sample asymptotic prediction optimality, estimation consistency, and asymptotic distribution of the PUMA estimator. Extensive simulations and a real-world application further demonstrate the empirical advantages of the proposed method over existing state-of-the-art approaches.
翻译:暂无翻译