High-resolution estimates of population health indicators are critical for precision public health. We propose a method for high-resolution estimation that fuses distinct data sources: an unbiased, low-resolution data source (e.g. aggregated administrative data) and a potentially biased, high-resolution data source (e.g. individual-level online survey responses). We assume that the potentially biased, high-resolution data source is generated from the population under a model of sampling bias where observables can have arbitrary impact on the probability of response but the difference in the log probabilities of response between units with the same observables is linear in the difference between sufficient statistics of their observables and outcomes. Our data fusion method learns a distribution that is closest (in the sense of KL divergence) to the online survey distribution and consistent with the aggregated administrative data and our model of sampling bias. This approach significantly reduces bias in high-resolution estimates compared to baselines that rely on a single data source alone on a testbed that includes repeated measurements of three indicators measured by both the (online) Household Pulse Survey and ground-truth data sources at two geographic resolutions over the same time period.
翻译:人口健康指标的高分辨率估计对于精准公共卫生至关重要。我们提出了一种融合不同数据源的高分辨率估计方法:无偏的低分辨率数据源(例如聚合行政数据)与可能存在偏倚的高分辨率数据源(例如个体层面在线调查响应)。我们假设存在潜在偏倚的高分辨率数据源是在抽样偏倚模型下从总体生成的,其中可观测变量对响应概率的影响是任意的,但具有相同可观测变量的单元之间响应概率对数差与其可观测变量和结果的充分统计量之差呈线性关系。我们的数据融合方法学习一个在KL散度意义上最接近在线调查分布的分布,该分布同时与聚合行政数据及我们的抽样偏倚模型保持一致。在包含同一时期两个地理分辨率下(在线)家庭脉搏调查与真实数据源重复测量的三项指标测试平台上,该方法相较于仅依赖单一数据源的基线显著降低了高分辨率估计中的偏倚。