Despite CLIP being the foundation model in numerous vision-language applications, the CLIP suffers from a severe text spotting bias. Such bias causes CLIP models to 'Parrot' the visual text embedded within images while disregarding the authentic visual semantics. We uncover that in the most popular image-text dataset LAION-2B, the captions also densely parrot (spell) the text embedded in images. Our analysis shows that around 50% of images are embedded with visual text content, and 90% of their captions more or less parrot the visual text. Based on such observation, we thoroughly inspect the different released versions of CLIP models and verify that the visual text is the dominant factor in measuring the LAION-style image-text similarity for these models. To examine whether these parrot captions shape the text spotting bias, we train a series of CLIP models with LAION subsets curated by different parrot-caption-oriented criteria. We show that training with parrot captions easily shapes such bias but harms the expected visual-language representation learning in CLIP models. This suggests that it is urgent to revisit either the design of CLIP-like models or the existing image-text dataset curation pipeline built on CLIP score filtering.
翻译:尽管CLIP是众多视觉-语言应用中的基础模型,但其存在严重的文字识别偏差。这种偏差导致CLIP模型会"鹦鹉学舌"般复述图像中的视觉文字,而忽略真实的视觉语义。我们发现,在最大规模的图像-文本数据集LAION-2B中,描述文本同样密集地"鹦鹉式"复述(拼写)了图像中的嵌入文字。分析显示,约50%的图像包含视觉文字内容,且其中90%的描述文本或多或少地复述了这些视觉文字。基于此观察,我们全面检测了CLIP模型的不同发布版本,验证了视觉文字是衡量这些模型在LAION风格图像-文本相似度的主导因素。为探究这些"鹦鹉式描述"是否塑造了文字识别偏差,我们使用基于不同"鹦鹉式描述"准则筛选的LAION子集训练了一系列CLIP模型。研究表明,使用"鹦鹉式描述"进行训练虽然容易形成此类偏差,但会损害CLIP模型期望的视觉-语言表征学习效果。这提示我们亟需重新审视CLIP类模型的设计方案,或基于CLIP分数过滤的现有图像-文本数据集构建流程。