Self-supervised learning (SSL) has emerged as a promising alternative to create supervisory signals to real-world problems, avoiding the extensive cost of manual labeling. SSL is particularly attractive for unsupervised tasks such as anomaly detection (AD), where labeled anomalies are rare or often nonexistent. A large catalog of augmentation functions has been used for SSL-based AD (SSAD) on image data, and recent works have reported that the type of augmentation has a significant impact on accuracy. Motivated by those, this work sets out to put image-based SSAD under a larger lens and investigate the role of data augmentation in SSAD. Through extensive experiments on 3 different detector models and across 420 AD tasks, we provide comprehensive numerical and visual evidences that the alignment between data augmentation and anomaly-generating mechanism is the key to the success of SSAD, and in the lack thereof, SSL may even impair accuracy. To the best of our knowledge, this is the first meta-analysis on the role of data augmentation in SSAD.
翻译:自监督学习(SSL)已成为一种有前景的方法,为现实问题创建监督信号,避免了人工标注的高昂成本。SSL对于异常检测(AD)等无监督任务尤其具有吸引力,因为此类任务中标注异常样本稀少甚至不存在。大量数据增强函数已被用于基于SSL的图像异常检测(SSAD),近期研究表明增强类型对精度有显著影响。受此启发,本研究旨在更全面地审视基于图像的SSAD方法,探究数据增强在SSAD中的作用。通过在3种不同检测器模型和420个AD任务上的广泛实验,我们提供了详尽的数值和视觉证据,证明数据增强与异常生成机制的对齐是SSAD成功的关键,缺乏这种对齐时,SSL甚至可能降低精度。据我们所知,这是首次针对数据增强在SSAD中的作用开展元分析研究。