Vision-Language Pre-training has demonstrated its remarkable zero-shot recognition ability and potential to learn generalizable visual representations from language supervision. Taking a step ahead, language-supervised semantic segmentation enables spatial localization of textual inputs by learning pixel grouping solely from image-text pairs. Nevertheless, the state-of-the-art suffers from clear semantic gaps between visual and textual modality: plenty of visual concepts appeared in images are missing in their paired captions. Such semantic misalignment circulates in pre-training, leading to inferior zero-shot performance in dense predictions due to insufficient visual concepts captured in textual representations. To close such semantic gap, we propose Concept Curation (CoCu), a pipeline that leverages CLIP to compensate for the missing semantics. For each image-text pair, we establish a concept archive that maintains potential visually-matched concepts with our proposed vision-driven expansion and text-to-vision-guided ranking. Relevant concepts can thus be identified via cluster-guided sampling and fed into pre-training, thereby bridging the gap between visual and textual semantics. Extensive experiments over a broad suite of 8 segmentation benchmarks show that CoCu achieves superb zero-shot transfer performance and greatly boosts language-supervised segmentation baseline by a large margin, suggesting the value of bridging semantic gap in pre-training data.
翻译:视觉-语言预训练已展现出其卓越的零样本识别能力以及从语言监督中学习可泛化视觉表征的潜力。在此基础上,语言监督的语义分割通过仅从图像-文本对中学习像素分组,实现了对文本输入的语义空间定位。然而,现有最先进方法在视觉与文本模态间存在明显的语义鸿沟:图像中出现的众多视觉概念在其配对标题中缺失。这种语义错位在预训练过程中循环累积,导致文本表征捕获的视觉概念不足,从而在密集预测任务中表现欠佳。为弥合这一语义鸿沟,我们提出概念策展(CoCu)流程,利用CLIP补偿缺失的语义信息。针对每个图像-文本对,我们构建一个概念档案库,通过所提的视觉驱动扩展和文本到视觉引导排序维护潜在的视觉匹配概念。进而通过聚类引导采样识别相关概念,并将其纳入预训练,从而桥接视觉与文本语义。在涵盖8个语义分割基准的广泛实验表明,CoCu实现了卓越的零样本迁移性能,并大幅提升了语言监督语义分割基线,突显了弥合预训练数据中语义鸿沟的重要价值。