When mapping subnational health and demographic indicators, direct weighted estimators of small area means based on household survey data can be unreliable when data are limited. If survey microdata are available, unit level models can relate individual survey responses to unit level auxiliary covariates and explicitly account for spatial dependence and between area variation using random effects. These models can produce estimators with improved precision, but often neglect to account for the design of the surveys used to collect data. Pseudo-Bayesian approaches incorporate sampling weights to address informative sampling when using such models to conduct population inference but credible sets based on the resulting pseudo-posterior distributions can be poorly calibrated without adjustment. We outline a pseudo-Bayesian strategy for small area estimation that addresses informative sampling and incorporates a post-processing rescaling step that produces credible sets with close to nominal empirical frequentist coverage rates. We compare our approach with existing design-based and model-based estimators using real and simulated data.
翻译:在绘制国家以下层面的健康与人口统计指标时,基于住户调查数据的小域均值直接加权估计量在数据有限的情况下可能不可靠。若调查微观数据可用,单元级模型可将个体调查响应与单元级辅助协变量相关联,并通过随机效应显式地考虑空间依赖性和区域间变异。这些模型能够产生精度更高的估计量,但常忽略用于收集数据的调查设计。伪贝叶斯方法通过纳入抽样权重来处理信息性抽样,从而在使用此类模型进行总体推断时加以调整,但基于所得伪后验分布的置信区间若不经校准可能覆盖性能欠佳。我们提出了一种用于小域估计的伪贝叶斯策略,该策略既处理信息性抽样,又引入后处理重缩放步骤,从而产生具有接近名义经验频率覆盖率的置信区间。我们利用真实数据和模拟数据,将我们的方法与现有基于设计和基于模型的估计量进行了比较。