Recent success of large-scale Contrastive Language-Image Pre-training (CLIP) has led to great promise in zero-shot semantic segmentation by transferring image-text aligned knowledge to pixel-level classification. However, existing methods usually require an additional image encoder or retraining/tuning the CLIP module. Here, we present a cost-effective strategy using text-prompt learning that keeps the entire CLIP module frozen while fully leveraging its rich information. Specifically, we propose a novel Zero-shot segmentation with Optimal Transport (ZegOT) method that matches multiple text prompts with frozen image embeddings through optimal transport, which allows each text prompt to efficiently focus on specific semantic attributes. Additionally, we propose Deep Local Feature Alignment (DLFA) that deeply aligns the text prompts with intermediate local feature of the frozen image encoder layers, which significantly boosts the zero-shot segmentation performance. Through extensive experiments on benchmark datasets, we show that our method achieves the state-of-the-art (SOTA) performance with only x7 lighter parameters compared to previous SOTA approaches.
翻译:大规模对比语言-图像预训练(CLIP)的成功通过将图像-文本对齐知识迁移至像素级分类,为零样本语义分割带来了巨大前景。然而,现有方法通常需要额外的图像编码器或对CLIP模块重新训练/微调。本文提出了一种成本高效的文本提示学习策略,该策略在完全冻结CLIP模块的同时充分利用其丰富信息。具体而言,我们提出了一种基于最优输运的零样本分割(ZegOT)方法,通过最优输运将多个文本提示与冻结的图像嵌入对齐,使每个文本提示能够高效聚焦于特定语义属性。此外,我们提出了深度局部特征对齐(DLFA),该机制将文本提示与冻结图像编码器层的中间局部特征进行深度对齐,显著提升了零样本分割性能。通过在基准数据集上的大量实验,我们证明所提方法在参数仅为先前最优方法1/7的条件下,实现了最优性能。