We consider the problem of integrating a small probability sample (ps) and a non-probability sample (nps). By definition, for the nps, there are no survey weights, but for the ps, there are survey weights. The key issue is that the nps, although much larger than the ps, can lead to a biased estimator of a finite population quantity but with much smaller variance. We begin with a relatively simple problem in which the population is assumed to be homogeneous and there are no common units in the ps and the nps. We assume that there are covariates and responses for everyone in the two samples, and there are no covariates available for the nonsampled units. We use the nps (ps) to construct a prior for the ps (nps). We also introduce partial discounting to avoid a dominance of the prior. We use Bayesian predictive inference for the finite population mean. In our illustrative example on body mass index and our simulation study, we compare the relative performance of alternative procedures and demonstrate that our procedure leads to improved estimates over the ps only estimate.
翻译:我们考虑将一个小型概率样本(ps)与一个非概率样本(nps)进行整合的问题。根据定义,非概率样本没有调查权重,而概率样本具有调查权重。关键问题在于,尽管非概率样本的规模远大于概率样本,但其可能对有限总体量的估计产生有偏结果,不过方差却小得多。我们首先处理一个相对简单的问题:假设总体是同质的,且概率样本与非概率样本之间没有重叠单元。我们假设两个样本中的所有个体均具有协变量和响应变量数据,但未抽样单元无可用协变量。我们利用非概率样本(概率样本)为概率样本(非概率样本)构建先验分布,同时引入部分折扣机制以避免先验主导。我们采用贝叶斯预测推断估计有限总体均值。在关于身体质量指数的示例分析和模拟研究中,我们比较了不同方法的相对表现,并证明我们的方法相比仅使用概率样本的估计结果具有更优的估计效果。