Time series are generated in diverse domains such as economic, traffic, health, and energy, where forecasting of future values has numerous important applications. Not surprisingly, many forecasting methods are being proposed. To ensure progress, it is essential to be able to study and compare such methods empirically in a comprehensive and reliable manner. To achieve this, we propose TFB, an automated benchmark for Time Series Forecasting (TSF) methods. TFB advances the state-of-the-art by addressing shortcomings related to datasets, comparison methods, and evaluation pipelines: 1) insufficient coverage of data domains, 2) stereotype bias against traditional methods, and 3) inconsistent and inflexible pipelines. To achieve better domain coverage, we include datasets from 10 different domains: traffic, electricity, energy, the environment, nature, economic, stock markets, banking, health, and the web. We also provide a time series characterization to ensure that the selected datasets are comprehensive. To remove biases against some methods, we include a diverse range of methods, including statistical learning, machine learning, and deep learning methods, and we also support a variety of evaluation strategies and metrics to ensure a more comprehensive evaluations of different methods. To support the integration of different methods into the benchmark and enable fair comparisons, TFB features a flexible and scalable pipeline that eliminates biases. Next, we employ TFB to perform a thorough evaluation of 21 Univariate Time Series Forecasting (UTSF) methods on 8,068 univariate time series and 14 Multivariate Time Series Forecasting (MTSF) methods on 25 datasets. The benchmark code and data are available at https://github.com/decisionintelligence/TFB.
翻译:时间序列广泛产生于经济、交通、健康、能源等多个领域,对未来值的预测具有众多重要应用。因此,不断有新的预测方法被提出。为了确保研究进展,必须能够以全面可靠的方式对这些方法进行实证研究和比较。为此,我们提出了TFB——一个面向时间序列预测方法的自动化基准测试框架。TFB通过解决数据集、比较方法和评估流程三方面的不足,推动了该领域的发展:1)数据领域覆盖不足;2)对传统方法的刻板偏见;3)评估流程不一致且缺乏灵活性。为提升领域覆盖度,我们纳入了来自10个不同领域的数集:交通、电力、能源、环境、自然、经济、股票市场、金融、健康及网络领域。我们还提供了时间序列特征描述,以确保所选数据集的全面性。为消除对某些方法的偏见,我们涵盖了包括统计学习、机器学习和深度学习方法在内的多种方法,并支持多种评估策略与指标,以实现对不同方法更全面的评估。为支持不同方法集成到基准测试中并实现公平比较,TFB采用了一个灵活、可扩展且能消除偏见的评估流程。基于此,我们使用TFB对21种单变量时间序列预测方法在8,068条单变量时间序列上,以及14种多变量时间序列预测方法在25个数据集上进行了全面评估。基准测试代码与数据已公开于https://github.com/decisionintelligence/TFB。