We uncover a surprising multilingual bias occurring in a popular class of multimodal vision-language models (VLMs). Including an image in the query to a LLaVA-style VLM significantly increases the likelihood of the model returning an English response, regardless of the language of the query. This paper investigates the causes of this loss with a two-pronged approach that combines extensive ablation of the design space with a mechanistic analysis of the models' internal representations of image and text inputs. Both approaches indicate that the issue stems in the language modelling component of the LLaVA model. Statistically, we find that switching the language backbone for a bilingual language model has the strongest effect on reducing this error. Mechanistically, we provide compelling evidence that visual inputs are not mapped to a similar space as text ones, and that intervening on intermediary attention layers can reduce this bias. Our findings provide important insights to researchers and engineers seeking to understand the crossover between multimodal and multilingual spaces, and contribute to the goal of developing capable and inclusive VLMs for non-English contexts.
翻译:我们揭示了一类流行的多模态视觉语言模型(VLMs)中存在的令人惊讶的多语言偏见。在向LLaVA风格的VLM查询中包含图像会显著增加模型返回英语回复的可能性,无论查询使用何种语言。本文通过结合设计空间的广泛消融研究与模型对图像和文本输入内部表征的机制分析,采用双管齐下的方法探究了这种损失的原因。两种方法均表明,问题源于LLaVA模型的语言建模组件。从统计上看,我们发现将语言主干网络切换为双语语言模型对减少此类错误效果最为显著。从机制上,我们提供了有力证据,表明视觉输入并未映射到与文本输入相似的空间,并且对中间注意力层进行干预可以减少这种偏见。我们的研究结果为寻求理解多模态与多语言空间交叉的研究人员和工程师提供了重要见解,并有助于为非英语语境开发能力强且包容的VLMs这一目标做出贡献。