Accurate weather forecasting is crucial in various sectors, impacting decision-making processes and societal events. Data-driven approaches based on machine learning models have recently emerged as a promising alternative to numerical weather prediction models given their potential to capture physics of different scales from historical data and the significantly lower computational cost during the prediction stage. Renowned for its state-of-the-art performance across diverse domains, the Transformer model has also gained popularity in machine learning weather prediction. Yet applying Transformer architectures to weather forecasting, particularly on a global scale is computationally challenging due to the quadratic complexity of attention and the quadratic increase in spatial points as resolution increases. In this work, we propose a factorized-attention-based model tailored for spherical geometries to mitigate this issue. More specifically, it utilizes multi-dimensional factorized kernels that convolve over different axes where the computational complexity of the kernel is only quadratic to the axial resolution instead of overall resolution. The deterministic forecasting accuracy of the proposed model on $1.5^\circ$ and 0-7 days' lead time is on par with state-of-the-art purely data-driven machine learning weather prediction models. We also showcase the proposed model holds great potential to push forward the Pareto front of accuracy-efficiency for Transformer weather models, where it can achieve better accuracy with less computational cost compared to Transformer based models with standard attention.
翻译:准确的气象预报在诸多领域具有关键作用,影响着决策制定与社会活动。近年来,基于机器学习模型的数据驱动方法因其能从历史数据中捕捉不同尺度的物理规律,以及预测阶段计算成本显著降低的优势,已成为数值天气预报模型的有力替代方案。Transformer模型以其在多个领域的先进性能著称,在机器学习气象预测中也逐渐流行。然而,将Transformer架构应用于气象预报(特别是全球尺度)仍面临计算挑战——这是由于注意力机制的二次复杂度以及空间点数量随分辨率提升呈平方级增长所致。为解决该问题,本文提出一种专为球面几何设计的分解注意力模型。具体而言,该模型采用多维分解核沿不同轴向进行卷积运算,其计算复杂度仅与轴向分辨率呈二次关系,而非整体分辨率。该模型在1.5°分辨率、0-7天预测时效下的确定性预报精度与现有最先进的纯数据驱动机器学习气象预测模型相当。我们还证明,该模型在推动Transformer气象模型精度-效率帕累托前沿方面具有巨大潜力——相较于采用标准注意力的Transformer模型,它能以更低计算成本实现更优精度。