Rare event prediction involves identifying and forecasting events with a low probability using machine learning and data analysis. Due to the imbalanced data distributions, where the frequency of common events vastly outweighs that of rare events, it requires using specialized methods within each step of the machine learning pipeline, i.e., from data processing to algorithms to evaluation protocols. Predicting the occurrences of rare events is important for real-world applications, such as Industry 4.0, and is an active research area in statistical and machine learning. This paper comprehensively reviews the current approaches for rare event prediction along four dimensions: rare event data, data processing, algorithmic approaches, and evaluation approaches. Specifically, we consider 73 datasets from different modalities (i.e., numerical, image, text, and audio), four major categories of data processing, five major algorithmic groupings, and two broader evaluation approaches. This paper aims to identify gaps in the current literature and highlight the challenges of predicting rare events. It also suggests potential research directions, which can help guide practitioners and researchers.
翻译:罕见事件预测涉及利用机器学习与数据分析识别和预测低概率事件。由于数据分布不平衡(常见事件频率远超罕见事件),需要在机器学习流程的每个步骤(从数据处理到算法到评估协议)中采用专门方法。预测罕见事件的发生对工业4.0等实际应用具有重要意义,且是统计学与机器学习领域的一个活跃研究方向。本文从四个维度全面综述当前罕见事件预测方法:罕见事件数据、数据处理、算法方法以及评估方法。具体而言,我们考虑了来自不同模态(即数值、图像、文本和音频)的73个数据集、四大类数据处理方法、五种主要算法分组以及两种更广泛的评估方法。本文旨在识别当前文献中的空白,突出预测罕见事件的挑战,并提出潜在的研究方向,以期为实践者和研究者提供指导。