Seasonal influenza causes on average 425,000 hospitalizations and 32,000 deaths per year in the United States. Forecasts of influenza-like illness (ILI) -- a surrogate for the proportion of patients infected with influenza -- support public health decision making. The goal of an ensemble forecast of ILI is to increase accuracy and calibration compared to individual forecasts and to provide a single, cohesive prediction of future influenza. However, an ensemble may be composed of models that produce similar forecasts, causing issues with ensemble forecast performance and non-identifiability. To improve upon the above issues we propose a novel Cluster-Aggregate-Pool or `CAP' ensemble algorithm that first clusters together individual forecasts, aggregates individual models that belong to the same cluster into a single forecast (called a cluster forecast), and then pools together cluster forecasts via a linear pool. When compared to a non-CAP approach, we find that a CAP ensemble improves calibration by approximately 10% while maintaining similar accuracy to non-CAP alternatives. In addition, our CAP algorithm (i) generalizes past ensemble work associated with influenza forecasting and introduces a framework for future ensemble work, (ii) automatically accounts for missing forecasts from individual models, (iii) allows public health officials to participate in the ensemble by assigning individual models to clusters, and (iv) provide an additional signal about when peak influenza may be near.
翻译:季节性流感在美国每年平均导致约42.5万例住院和3.2万例死亡。流感样疾病(ILI)——作为流感患者比例的代表指标——的预测有助于公共卫生决策。ILI集成预测的目标是相较于单一预测模型提高准确性和校准度,并为未来流感趋势提供统一且连贯的预测。然而,集成模型可能包含产生相似预测结果的不同模型,这会导致集成预测性能下降及模型不可识别性问题。为解决上述问题,我们提出一种新颖的聚类-聚合-池化(CAP)集成算法:首先将独立预测结果进行聚类,将属于同一聚类的单个模型聚合为单一预测(称为聚类预测),再通过线性池化方法整合所有聚类预测。与非CAP方法相比,CAP集成在校准度上提升约10%,同时保持与非CAP方法相当的准确性。此外,CAP算法:(i)拓展了流感预测领域既往的集成研究工作,并构建了未来集成研究的框架;(ii)自动处理单个模型的缺失预测问题;(iii)允许公共卫生官员通过将单个模型分配至特定聚类参与集成;(iv)提供流感高峰临近的额外预警信号。