If part of a population is hidden but two or more sources are available that each cover parts of this population, dual- or multiple-system(s) estimation can be applied to estimate this population. For this it is common to use the log-linear model, estimated with maximum likelihood. These maximum likelihood estimates are based on a non-linear model and therefore suffer from finite-sample bias, which can be substantial in case of small samples or a small population size. This problem was recognised by Chapman, who derived an estimator with good small sample properties in case of two available sources. However, he did not derive an estimator for more than two sources. We propose an estimator that is an extension of Chapman's estimator to three or more sources and compare this estimator with other bias-reduced estimators in a simulation study. The proposed estimator performs well, and much better than the other estimators. A real data example on homelessness in the Netherlands shows that our proposed model can make a substantial difference.
翻译:如果总体的一部分是隐藏的,但存在两个或多个覆盖该部分总体的来源,则可以采用双系统或多系统估计来估计该总体。通常,我们使用对数线性模型,并通过最大似然法进行估计。这些最大似然估计基于非线性模型,因此存在有限样本偏差,在样本量较小或总体规模较小时,偏差可能相当显著。这一问题的识别归功于查普曼,他推导出了一种在只有两个可用来源的情况下具有良好小样本性质的估计量。然而,他并未推导出适用于两个以上来源的估计量。我们提出了一种将查普曼估计量扩展至三个或更多来源的估计量,并通过模拟研究将其与其他偏差缩减估计量进行了比较。我们提出的估计量表现良好,且远优于其他估计量。关于荷兰无家可归者的真实数据实例表明,我们所提出的模型可能带来显著差异。