The network scale-up method (NSUM) is a cost-effective approach to estimating the size or prevalence of a group of people that is hard to reach through a standard survey. The basic NSUM involves two steps: estimating respondents' degrees by one of various methods (in this paper we focus on the probe group method which uses the number of people a respondent knows in various groups of known size), and estimating the prevalence of the hard-to-reach population of interest using respondents' estimated degrees and the number of people they report knowing in the hard-to-reach group. Each of these two steps involves taking either an average of ratios or a ratio of averages. Using the ratio of averages for each step has so far been the most common approach. However, we present theoretical arguments that using the average of ratios at the second, prevalence-estimation step often has lower mean squared error when the random mixing assumption is violated, which seems likely in practice; this estimator which uses the ratio of averages for degree estimates and the average of ratios for prevalence was proposed early in NSUM development but has largely been unexplored and unused. Simulation results using an example network data set also support these findings. Based on this theoretical and empirical evidence, we suggest that future surveys that use a simple estimator may want to use this mixed estimator, and estimation methods based on this estimator may produce new improvements.
翻译:网络规模法(NSUM)是一种经济高效的方法,用于估计难以通过标准调查触及的人群规模或流行率。基本NSUM包含两个步骤:首先通过多种方法估算受访者的人际网络规模(本文聚焦于探测组方法,该方法利用受访者在已知规模的群体中认识的人数来估算),然后利用受访者的人际网络规模及其报告在目标难触及群体中认识的人数,估算该群体的流行率。这两个步骤均涉及比率均值与均值比率的选择。以往研究中,各步骤均采用均值比率最为常见。然而,我们通过理论论证表明,当随机混合假设(实践中常被违反)不成立时,在第二步流行率估算中采用比率均值通常能获得更低的均方误差。这种在人际网络规模估算中使用均值比率、在流行率估算中使用比率均值的混合估计量,虽在NSUM发展早期已被提出,却长期未被充分探索与应用。基于网络数据集的模拟结果也支持这一发现。综合理论与实证证据,我们建议未来使用简单估计量的调查可优先采用此混合估计量,而在此基础上发展的估计方法有望实现新的改进。