Time-series modeling in process industries faces the challenge of dealing with complex, multi-faceted, and evolving data characteristics. Conventional single model approaches often struggle to capture the interplay of diverse dynamics, resulting in suboptimal forecasts. Addressing this, we introduce the Recency-Weighted Temporally-Segmented (ReWTS, pronounced `roots') ensemble model, a novel chunk-based approach for multi-step forecasting. The key characteristics of the ReWTS model are twofold: 1) It facilitates specialization of models into different dynamics by segmenting the training data into `chunks' of data and training one model per chunk. 2) During inference, an optimization procedure assesses each model on the recent past and selects the active models, such that the appropriate mixture of previously learned dynamics can be recalled to forecast the future. This method not only captures the nuances of each period, but also adapts more effectively to changes over time compared to conventional `global' models trained on all data in one go. We present a comparative analysis, utilizing two years of data from a wastewater treatment plant and a drinking water treatment plant in Norway, demonstrating the ReWTS ensemble's superiority. It consistently outperforms the global model in terms of mean squared forecasting error across various model architectures by 10-70\% on both datasets, notably exhibiting greater resilience to outliers. This approach shows promise in developing automatic, adaptable forecasting models for decision-making and control systems in process industries and other complex systems.
翻译:过程工业中的时间序列建模面临着处理复杂、多面且不断演变的数据特征的挑战。传统单一模型方法往往难以捕捉多种动态特性的相互作用,导致预测效果欠佳。针对这一问题,我们提出了基于近期加权的时序分段集成模型(ReWTS,发音为'roots'),这是一种新颖的基于数据块的多步预测方法。ReWTS模型的关键特征有两方面:1) 通过将训练数据分割成多个数据块,并为每个数据块训练一个模型,促使模型专门化处理不同的动态特性。2) 在推理过程中,通过优化程序评估每个模型在近期数据上的表现,并选择活跃模型,从而能够调用先前学习到的适当动态组合来预测未来。与一次性在所有数据上训练的传统"全局"模型相比,该方法不仅能够捕捉每个时间段的细微特征,还能更有效地适应随时间发生的变化。我们利用挪威某污水处理厂和饮用水处理厂两年的数据进行了对比分析,证明了ReWTS集成模型的优越性。在两个数据集上,该模型在不同模型架构下的均方预测误差始终优于全局模型10%至70%,并且表现出更强的异常值鲁棒性。该方法为过程工业及其他复杂系统中的决策和控制系统开发自动、自适应的预测模型展现了良好前景。