Contextual features are important data sources for building citywide crowd mobility prediction models. However, the difficulty of applying context lies in the unknown generalizability of contextual features (e.g., weather, holiday, and points of interests) and context modeling techniques across different scenarios. In this paper, we present a unified analytic framework and a large-scale benchmark for evaluating context generalizability. The benchmark includes crowd mobility data, contextual data, and advanced prediction models. We conduct comprehensive experiments in several crowd mobility prediction tasks such as bike flow, metro passenger flow, and electric vehicle charging demand. Our results reveal several important observations: (1) Using more contextual features may not always result in better prediction with existing context modeling techniques; in particular, the combination of holiday and temporal position can provide more generalizable beneficial information than other contextual feature combinations. (2) In context modeling techniques, using a gated unit to incorporate raw contextual features into the deep prediction model has good generalizability. Besides, we offer several suggestions about incorporating contextual factors for building crowd mobility prediction applications. From our findings, we call for future research efforts devoted to developing new context modeling solutions.
翻译:上下文特征是构建城市级人群流动性预测模型的重要数据来源。然而,应用上下文的难点在于上下文特征(例如天气、节假日和兴趣点)及上下文建模技术在不同场景下的可泛化性未知。本文提出了一个统一的分析框架和一个大规模基准,用于评估上下文可泛化性。该基准包含人群流动性数据、上下文数据以及先进的预测模型。我们在多个城市人群流动性预测任务中(如自行车流量、地铁客流量及电动汽车充电需求)开展了全面实验。研究结果揭示了若干重要发现:(1)利用更多的上下文特征并不总是能借助现有上下文建模技术带来更优的预测效果;具体而言,节假日与时间位置的组合能比其他上下文特征组合提供更具泛化性的有益信息。(2)在上下文建模技术中,使用门控单元将原始上下文特征融入深度预测模型具有良好的泛化性。此外,我们就如何纳入上下文因子构建人群流动性预测应用提出了若干建议。基于以上发现,我们呼吁未来研究致力于开发新的上下文建模解决方案。