Internet memes represent a popular form of multimodal online communication and often use figurative elements to convey layered meaning through the combination of text and images. However, it remains largely unclear how multimodal large language models (MLLMs) combine and interpret visual and textual information to identify figurative meaning in memes. To address this gap, we evaluate eight state-of-the-art generative MLLMs across three datasets on their ability to detect and explain six types of figurative meaning. In addition, we conduct a human evaluation of the explanations generated by these MLLMs, assessing whether the provided reasoning supports the predicted label and whether it remains faithful to the original meme content. Our findings indicate that all models exhibit a strong bias to associate a meme with figurative meaning, even when no such meaning is present. Qualitative analysis further shows that correct predictions are not always accompanied by faithful explanations.
翻译:互联网模因是一种流行的多模态在线交流形式,常通过文本与图像结合使用比喻元素来传达多层次意义。然而,多模态大语言模型如何结合并理解视觉与文本信息以识别模因中的比喻意义,目前仍不清楚。为填补这一空白,我们基于三个数据集评估了八种最先进的生成式多模态大语言模型,考察其检测和解释六类比喻意义的能力。此外,我们还对这些模型生成的解释进行了人工评估,检查所提供推理是否支持预测标签,以及其是否忠实于原始模因内容。研究结果表明,所有模型均表现出将模因与比喻意义关联的强烈偏差,即使当模因中不存在此类意义时也是如此。定性分析进一步显示,正确的预测并不总是伴随着忠实的解释。