It is important to predict how the Global Mean Temperature (GMT) will evolve in the next few decades. The ability to predict historical data is a necessary first step toward the actual goal of making long-range forecasts. This paper examines the advantage of statistical and simpler Machine Learning (ML) methods instead of directly using complex ML algorithms and Deep Learning Neural Networks (DNN). Often neglected data transformation methods prior to applying different algorithms have been used as a means of improving predictive accuracy. The GMT time series is treated both as a univariate time series and also cast as a regression problem. Some steps of data transformations were found to be effective. Various simple ML methods did as well or better than the more well-known ones showing merit in trying a large bouquet of algorithms as a first step. Fifty-six algorithms were subject to Box-Cox, Yeo-Johnson, and first-order differencing and compared with the absence of them. Predictions for the annual GMT testing data were better than that published so far, with the lowest RMSE value of 0.02 $^\circ$C. RMSE for five-year mean GMT values for the test data ranged from 0.00002 to 0.00036 $^\circ$C.
翻译:准确预测未来几十年全球平均气温(GMT)的演变趋势至关重要。对历史数据的预测能力是实现长期预报这一实际目标的必要前提。本文研究了统计方法与简单机器学习(ML)方法相较于直接使用复杂ML算法和深度学习神经网络(DNN)的优势。采用先前常被忽视的数据变换方法作为预处理手段,在应用不同算法前提升预测精度。我们将GMT时间序列同时作为单变量时间序列和回归问题进行建模。研究发现某些数据变换步骤具有显著效果。多种简单ML方法的表现达到甚至优于更著名的算法,这表明作为首要步骤尝试大量算法具有实际价值。对56种算法分别进行Box-Cox变换、Yeo-Johnson变换和一阶差分处理,并与未做变换的情况进行对比。年度GMT测试数据的预测结果优于现有文献,最低均方根误差(RMSE)达到0.02°C。五年均值GMT的测试数据RMSE介于0.00002至0.00036°C之间。