Subseasonal forecasting of the weather two to six weeks in advance is critical for resource allocation and advance disaster notice but poses many challenges for the forecasting community. At this forecast horizon, physics-based dynamical models have limited skill, and the targets for prediction depend in a complex manner on both local weather variables and global climate variables. Recently, machine learning methods have shown promise in advancing the state of the art but only at the cost of complex data curation, integrating expert knowledge with aggregation across multiple relevant data sources, file formats, and temporal and spatial resolutions. To streamline this process and accelerate future development, we introduce SubseasonalClimateUSA, a curated dataset for training and benchmarking subseasonal forecasting models in the United States. We use this dataset to benchmark a diverse suite of models, including operational dynamical models, classical meteorological baselines, and ten state-of-the-art machine learning and deep learning-based methods from the literature. Overall, our benchmarks suggest simple and effective ways to extend the accuracy of current operational models. SubseasonalClimateUSA is regularly updated and accessible via the https://github.com/microsoft/subseasonal_data/ Python package.
翻译:次季节天气预报(提前2至6周)对资源调配和灾害预警至关重要,但给预报领域带来诸多挑战。在此预报时间尺度上,基于物理的动力学模型预报能力有限,预测目标既受当地天气变量影响,又与全球气候变量存在复杂依赖关系。近年来,机器学习方法在提升预报水平方面展现出潜力——这些方法需要整合专家知识与多个数据源、不同文件格式及多样化时空分辨率的多源数据,但其复杂的数据处理流程限制了发展。为优化这一过程并加速未来研究,我们提出SubseasonalClimateUSA——一个用于美国次季节预报模型训练与基准测试的精选数据集。利用该数据集,我们对多种模型进行基准测试,包括业务动力学模型、经典气象基线模型以及文献中十种最先进的机器学习和深度学习方法。综合基准测试结果表明,存在简单有效的方式可提升当前业务模型的预报精度。该数据集定期更新,可通过https://github.com/microsoft/subseasonal_data/ Python工具包访问。