Deep learning for tabular data has garnered increasing attention in recent years, yet employing deep models for structured data remains challenging. While these models excel with unstructured data, their efficacy with structured data has been limited. Recent research has introduced retrieval-augmented models to address this gap, demonstrating promising results in supervised tasks such as classification and regression. In this work, we investigate using retrieval-augmented models for anomaly detection on tabular data. We propose a reconstruction-based approach in which a transformer model learns to reconstruct masked features of \textit{normal} samples. We test the effectiveness of KNN-based and attention-based modules to select relevant samples to help in the reconstruction process of the target sample. Our experiments on a benchmark of 31 tabular datasets reveal that augmenting this reconstruction-based anomaly detection (AD) method with non-parametric relationships via retrieval modules may significantly boost performance.
翻译:近年来,深度学习在表格数据上的应用日益受到关注,但将深度模型用于结构化数据仍面临挑战。尽管这些模型在非结构化数据上表现出色,但它们在结构化数据上的效能有限。近期研究引入了检索增强模型以弥合这一差距,并在分类、回归等监督任务中展现出良好前景。本文探索使用检索增强模型进行表格数据上的异常检测。我们提出一种基于重建的方法,其中Transformer模型学习重构“正常”样本的掩码特征。我们测试了基于KNN和基于注意力机制的模块在目标样本重建过程中选择相关样本的有效性。在31个表格数据集的基准实验表明,通过检索模块引入非参数关系来增强这种基于重建的异常检测方法,可显著提升性能。