Anomaly detection is crucial in various domains, such as finance, healthcare, and cybersecurity. In this paper, we propose a novel deep anomaly detection method for tabular data that leverages Non-Parametric Transformers (NPTs), a model initially proposed for supervised tasks, to capture both feature-feature and sample-sample dependencies. In a reconstruction-based framework, we train the NPT to reconstruct masked features of normal samples. In a non-parametric fashion, we leverage the whole training set during inference and use the model's ability to reconstruct the masked features during to generate an anomaly score. To the best of our knowledge, our proposed method is the first to successfully combine feature-feature and sample-sample dependencies for anomaly detection on tabular datasets. We evaluate our method on an extensive benchmark of 31 tabular datasets and demonstrate that our approach outperforms existing state-of-the-art methods based on the F1-score and AUROC by a significant margin.
翻译:异常检测在金融、医疗保健和网络安全等各个领域都至关重要。本文提出一种新颖的表格数据深度异常检测方法,该方法利用最初为监督任务设计的非参数Transformer(NPTs)来捕获特征-特征和样本-样本之间的依赖关系。在基于重构的框架中,我们训练NPT重构正常样本的掩码特征。以非参数方式,我们在推理过程中利用整个训练集,并利用模型重构掩码特征的能力来生成异常分数。据我们所知,我们提出的方法是首个成功结合特征-特征和样本-样本依赖关系以进行表格数据集异常检测的方法。我们在包含31个表格数据集的广泛基准上进行评估,并证明我们的方法基于F1分数和AUROC指标显著优于现有最先进方法。