Computer vision often treats perception as objective, and this assumption gets reflected in the way that datasets are collected and models are trained. For instance, image descriptions in different languages are typically assumed to be translations of the same semantic content. However, work in cross-cultural psychology and linguistics has shown that individuals differ in their visual perception depending on their cultural background and the language they speak. In this paper, we demonstrate significant differences in semantic content across languages in both dataset and model-produced captions. When data is multilingual as opposed to monolingual, captions have higher semantic coverage on average, as measured by scene graph, embedding, and linguistic complexity. For example, multilingual captions have on average 21.8% more objects, 24.5% more relations, and 27.1% more attributes than a set of monolingual captions. Moreover, models trained on content from different languages perform best against test data from those languages, while those trained on multilingual content perform consistently well across all evaluation data compositions. Our research provides implications for how diverse modes of perception can improve image understanding.
翻译:计算机视觉通常将感知视为客观的,这一假设体现在数据集收集和模型训练的方式中。例如,不同语言的图像描述通常被假定为对相同语义内容的翻译。然而,跨文化心理学和语言学的研究表明,个体的视觉感知因其文化背景和所用语言的不同而存在显著差异。在本文中,我们证明,无论是在数据集还是模型生成的描述中,不同语言间的语义内容都存在显著差异。与单语言描述相比,多语言描述在语义覆盖上平均更高,这一点通过场景图、嵌入和语言复杂度等指标得以衡量。例如,多语言描述平均比单语言描述集多包含21.8%的对象、24.5%的关系和27.1%的属性。此外,基于不同语言内容训练的模型在面对相应语言的测试数据时表现最佳,而基于多语言内容训练的模型在所有评估数据组合中均表现稳定。我们的研究揭示了多样化的感知模式如何能够改善图像理解。