GPS mobility data are increasingly used in epidemic modeling, allowing the construction of co-location networks or population flows. These trajectories typically exhibit high temporal sparsity because data collection is opportunistic and tied to phone use. Despite growing awareness of this limitation, the analysis and treatment of biases derived from it have been largely overlooked in existing epidemic modeling studies, raising concerns about the robustness of downstream inferences. We introduce a principled framework to quantify the impact of trajectory sparsity on key epidemic modeling outcomes across different levels of missingness. Our approach leverages a highly-complete dataset that exhibits both near-complete and sparse GPS trajectories. Near-complete trajectories provide baseline epidemic outcomes, while sparse trajectories provide realistic missingness patterns that we impose on the baseline to measure bias. In this way, we show how missing records can result in substantial underestimation of key measures of epidemic intensity, explained not only by the amount of missing data, but by more complex features of data missingness that should be taken into account when designing correction methods. Finally, we propose and evaluate a correction based on inverse probability weighting of network edges before epidemic model calibration, which is shown to reduce bias and parameter misspecification. We also demonstrate this correction on a separate anonymized sample from a commercial GPS mobility dataset and report on its effect. Together, our findings provide a first rigorous quantification of trajectory-sparsity bias in epidemic modeling, offering initial guidance on the treatment of this issue.
翻译:全球定位系统(GPS)移动数据在流行病建模中日益普及,可用于构建共位网络或人口流动模型。这类轨迹数据因采集具有机会性特征且与手机使用行为紧密关联,通常呈现高度时间稀疏性。尽管学界已逐渐认识到这一局限性,但现有流行病建模研究大多忽视由此产生的偏差分析与处理方法,这引发了对下游推断结果稳健性的担忧。我们提出了一套规范化的分析框架,旨在量化不同缺失程度下轨迹稀疏性对关键流行病建模结果的影响。该方法借助一个具备近完整与稀疏两种GPS轨迹特征的高完整性数据集:近完整轨迹提供基准流行病结果,而稀疏轨迹则提供现实缺失模式——我们将此模式施加于基准数据以测量偏差。通过这种方式,我们揭示了数据缺失如何导致流行病强度关键指标的显著低估,其成因不仅包括数据缺失量,更涉及数据缺失的复杂特征,这些特征在设计校正方法时应当纳入考量。最后,我们提出并评估了一种基于逆概率加权网络边的校正方法(在流行病模型校准前实施),实证表明该方法能有效减少偏差与参数误设。我们还在独立匿名的商业GPS移动数据集样本上验证了该校正方法的效果。综合而言,本研究首次对流行病建模中的轨迹稀疏性偏差进行了严格量化,为处理该问题提供了初步指导。