Although the mapping between sound and meaning in human language is assumed to be largely arbitrary, research in cognitive science has shown that there are non-trivial correlations between particular sounds and meanings across languages and demographic groups, a phenomenon known as sound symbolism. Among the many dimensions of meaning, sound symbolism is particularly salient and well-demonstrated with regards to cross-modal associations between language and the visual domain. In this work, we address the question of whether sound symbolism is reflected in vision-and-language models such as CLIP and Stable Diffusion. Using zero-shot knowledge probing to investigate the inherent knowledge of these models, we find strong evidence that they do show this pattern, paralleling the well-known kiki-bouba effect in psycholinguistics. Our work provides a novel method for demonstrating sound symbolism and understanding its nature using computational tools. Our code will be made publicly available.
翻译:尽管人类语言中声音与意义的映射被认为大体上是任意的,但认知科学研究表明,跨语言和跨人群的特定声音与意义之间存在非平凡的相关性,这种现象被称为声音象征。在意义的众多维度中,声音象征在语言与视觉领域的跨模态关联方面尤为显著且经过充分验证。本研究探讨视觉语言模型(如CLIP和Stable Diffusion)是否反映了声音象征。通过使用零样本知识探测来探究这些模型的内在知识,我们发现了强有力的证据表明它们确实呈现这种模式,与心理语言学中著名的kiki-bouba效应相呼应。我们的工作提供了一种利用计算工具展示声音象征并理解其本质的新方法。我们的代码将公开提供。