Anomaly detection (AD) tasks have been solved using machine learning algorithms in various domains and applications. The great majority of these algorithms use normal data to train a residual-based model and assign anomaly scores to unseen samples based on their dissimilarity with the learned normal regime. The underlying assumption of these approaches is that anomaly-free data is available for training. This is, however, often not the case in real-world operational settings, where the training data may be contaminated with an unknown fraction of abnormal samples. Training with contaminated data, in turn, inevitably leads to a deteriorated AD performance of the residual-based algorithms. In this paper we introduce a framework for a fully unsupervised refinement of contaminated training data for AD tasks. The framework is generic and can be applied to any residual-based machine learning model. We demonstrate the application of the framework to two public datasets of multivariate time series machine data from different application fields. We show its clear superiority over the naive approach of training with contaminated data without refinement. Moreover, we compare it to the ideal, unrealistic reference in which anomaly-free data would be available for training. The method is based on evaluating the contribution of individual samples to the generalization ability of a given model, and contrasting the contribution of anomalies with the one of normal samples. As a result, the proposed approach is comparable to, and often outperforms training with normal samples only.
翻译:异常检测(AD)任务已在多个领域和应用中通过机器学习算法得到解决。绝大多数这类算法使用正常数据训练基于残差的模型,并根据测试样本与学习到的正常模式的差异程度为其分配异常分数。这些方法的潜在假设是训练数据中不含异常样本。然而,在实际操作场景中,训练数据往往会被未知比例的异常样本污染,导致基于残差算法的异常检测性能不可避免地下降。本文提出了一种用于异常检测任务中受污染训练数据全无监督精炼的通用框架。该框架具有通用性,可应用于任意基于残差的机器学习模型。我们通过来自不同应用领域的两个多元时间序列机器数据集展示了该框架的应用效果。结果表明,该方法显著优于未精炼直接使用受污染数据训练的朴素方法。此外,我们将其与理想化的非现实参考基准(即使用纯正常数据进行训练)进行了比较。该方法通过评估单个样本对给定模型泛化能力的贡献,并对比异常样本与正常样本的贡献差异实现。最终,所提出方法的性能可与仅使用正常样本训练相媲美,且通常更优。