In this paper we study predictive mean matching mass imputation estimators to integrate data from probability and non-probability samples. We consider two approaches: matching predicted to observed ($\hat{y}-y$ matching) or predicted to predicted ($\hat{y}-\hat{y}$ matching) values. We prove the consistency of two semi-parametric mass imputation estimators based on these approaches and derive their variance and estimators of variance. Our approach can be employed with non-parametric regression techniques, such as kernel regression, and the analytical expression for variance can also be applied in nearest neighbour matching for non-probability samples. We conduct extensive simulation studies in order to compare the properties of this estimator with existing approaches, discuss the selection of $k$-nearest neighbours, and study the effects of model mis-specification. The paper finishes with empirical study in integration of job vacancy survey and vacancies submitted to public employment offices (admin and online data). Open source software is available for the proposed approaches.
翻译:本文研究利用预测均值匹配多重插补估计量整合概率样本与非概率样本数据的方法。我们考虑两种匹配策略:预测值与观测值匹配($\hat{y}-y$匹配)与预测值间匹配($\hat{y}-\hat{y}$匹配)。基于这两种方法,我们证明了两种半参数多重插补估计量的一致性,推导出相应的方差及方差估计量。该方法可结合非参数回归技术(如核回归)使用,其方差解析表达式同样适用于非概率样本的最近邻匹配。通过大量模拟研究,我们比较了该估计量与现有方法的性质,讨论了$k$近邻选择问题,并分析模型设定错误的效应。最后,本文通过整合职位空缺调查数据与公共就业机构提交的职位信息(行政数据与在线数据)开展实证研究。所提方法对应的开源软件已公开可用。