Anomaly detection is vital in many domains, such as finance, healthcare, and cybersecurity. In this paper, we propose a novel deep anomaly detection method for tabular data that leverages Non-Parametric Transformers (NPTs), a model initially proposed for supervised tasks, to capture both feature-feature and sample-sample dependencies. In a reconstruction-based framework, we train the NPT to reconstruct masked features of normal samples. In a non-parametric fashion, we leverage the whole training set during inference and use the model's ability to reconstruct the masked features to generate an anomaly score. To the best of our knowledge, this is the first work to successfully combine feature-feature and sample-sample dependencies for anomaly detection on tabular datasets. Through extensive experiments on 31 benchmark tabular datasets, we demonstrate that our method achieves state-of-the-art performance, outperforming existing methods by 1.7% and 1.2% in terms of F1-score and AUROC, respectively. Our ablation study provides evidence that modeling both types of dependencies is crucial for anomaly detection on tabular data.
翻译:异常检测在金融、医疗和网络安全等诸多领域至关重要。本文提出一种新型的表格数据深度异常检测方法,该方法利用最初为监督任务设计的非参数Transformer(NPT)模型,同时捕捉特征-特征和样本-样本依赖性。在基于重构的框架中,我们训练NPT以重构正常样本的掩码特征。采用非参数方式,我们在推理过程中利用整个训练集,通过模型重构掩码特征的能力生成异常评分。据我们所知,这是首个成功结合特征-特征和样本-样本依赖性进行表格数据集异常检测的工作。通过在31个基准表格数据集上的大量实验,我们证明所提方法在F1分数和AUROC指标上分别以1.7%和1.2%的优势超越现有方法,达到最先进性能。消融研究提供的证据表明,对两类依赖性的建模对于表格数据的异常检测至关重要。