Visual anomaly classification and segmentation are vital for automating industrial quality inspection. The focus of prior research in the field has been on training custom models for each quality inspection task, which requires task-specific images and annotation. In this paper we move away from this regime, addressing zero-shot and few-normal-shot anomaly classification and segmentation. Recently CLIP, a vision-language model, has shown revolutionary generality with competitive zero-/few-shot performance in comparison to full-supervision. But CLIP falls short on anomaly classification and segmentation tasks. Hence, we propose window-based CLIP (WinCLIP) with (1) a compositional ensemble on state words and prompt templates and (2) efficient extraction and aggregation of window/patch/image-level features aligned with text. We also propose its few-normal-shot extension WinCLIP+, which uses complementary information from normal images. In MVTec-AD (and VisA), without further tuning, WinCLIP achieves 91.8%/85.1% (78.1%/79.6%) AUROC in zero-shot anomaly classification and segmentation while WinCLIP+ does 93.1%/95.2% (83.8%/96.4%) in 1-normal-shot, surpassing state-of-the-art by large margins.
翻译:视觉异常分类与分割对于自动化工业质量检测至关重要。该领域先前的研究重点是为每项质量检测任务训练定制模型,这需要特定任务的图像和标注。本文突破这一范式,针对零样本和少正常样本下的异常分类与分割问题展开研究。近期视觉语言模型CLIP展现出革命性的泛化能力,在零样本/少样本任务中取得与全监督方法相媲美的表现。但CLIP在异常分类与分割任务中存在不足。为此,我们提出基于窗口的CLIP(WinCLIP),其包含:(1)基于状态词与提示模板的复合集成方法;(2)与文本对齐的窗口/块/图像级特征的高效提取与聚合。此外,我们提出其少正常样本扩展WinCLIP+,该方法利用正常图像中的互补信息。在MVTec-AD(及VisA)数据集上,无需额外微调,WinCLIP在零样本异常分类与分割任务中分别达到91.8%/85.1%(78.1%/79.6%)的AUROC值,而WinCLIP+在1个正常样本条件下分别达到93.1%/95.2%(83.8%/96.4%),显著超越现有最优方法。