The proliferation of social media has given rise to a new form of communication: memes. Memes are multimodal and often contain a combination of text and visual elements that convey meaning, humor, and cultural significance. While meme analysis has been an active area of research, little work has been done on unsupervised multimodal topic modeling of memes, which is important for content moderation, social media analysis, and cultural studies. We propose \textsf{PromptMTopic}, a novel multimodal prompt-based model designed to learn topics from both text and visual modalities by leveraging the language modeling capabilities of large language models. Our model effectively extracts and clusters topics learned from memes, considering the semantic interaction between the text and visual modalities. We evaluate our proposed model through extensive experiments on three real-world meme datasets, which demonstrate its superiority over state-of-the-art topic modeling baselines in learning descriptive topics in memes. Additionally, our qualitative analysis shows that \textsf{PromptMTopic} can identify meaningful and culturally relevant topics from memes. Our work contributes to the understanding of the topics and themes of memes, a crucial form of communication in today's society.\\ \red{\textbf{Disclaimer: This paper contains sensitive content that may be disturbing to some readers.}}
翻译:社交媒体的普及催生了一种新的传播形式:模因。模因是多模态的,通常包含文本与视觉元素的组合,用以传达意义、幽默及文化内涵。尽管模因分析已成为活跃的研究领域,但在无监督多模态模因主题建模方面尚缺乏深入研究,而这对于内容审核、社交媒体分析及文化研究至关重要。我们提出\textsf{PromptMTopic},一种新颖的多模态提示模型,通过利用大型语言模型的建模能力,从文本和视觉模态中学习主题。该模型有效提取并聚类模因中的学习主题,同时考虑文本与视觉模态间的语义交互。通过在三个真实世界模因数据集上进行广泛实验,我们验证了所提模型的优越性,其在学习模因描述性主题方面超越了当前最先进的主题建模基线。此外,定性分析表明,\textsf{PromptMTopic}能够识别模因中具有意义且与文化相关的主题。我们的工作有助于理解模因(当代社会中一种关键传播形式)的主题与主旨。\\ \red{\textbf{免责声明:本文包含可能令部分读者不适的敏感内容。}}