The paper explores the industrial multimodal Anomaly Detection (AD) task, which exploits point clouds and RGB images to localize anomalies. We introduce a novel light and fast framework that learns to map features from one modality to the other on nominal samples. At test time, anomalies are detected by pinpointing inconsistencies between observed and mapped features. Extensive experiments show that our approach achieves state-of-the-art detection and segmentation performance in both the standard and few-shot settings on the MVTec 3D-AD dataset while achieving faster inference and occupying less memory than previous multimodal AD methods. Moreover, we propose a layer-pruning technique to improve memory and time efficiency with a marginal sacrifice in performance.
翻译:本文探讨了工业多模态异常检测(AD)任务,该任务利用点云和RGB图像定位异常。我们提出了一种轻量级高速框架,通过学习在正常样本上将一种模态的特征映射到另一种模态。在测试时,通过识别观测特征与映射特征之间的不一致性来检测异常。大量实验表明,我们的方法在MVTec 3D-AD数据集的标准和少样本设置下均实现了最先进的检测与分割性能,同时推理速度更快且内存占用低于以往多模态AD方法。此外,我们提出了一种层剪枝技术,在性能轻微牺牲的情况下提升内存与时间效率。