Memes are the new-age conveyance mechanism for humor on social media sites. Memes often include an image and some text. Memes can be used to promote disinformation or hatred, thus it is crucial to investigate in details. We introduce Memotion 3, a new dataset with 10,000 annotated memes. Unlike other prevalent datasets in the domain, including prior iterations of Memotion, Memotion 3 introduces Hindi-English Codemixed memes while prior works in the area were limited to only the English memes. We describe the Memotion task, the data collection and the dataset creation methodologies. We also provide a baseline for the task. The baseline code and dataset will be made available at https://github.com/Shreyashm16/Memotion-3.0
翻译:梗图是社交媒体上新型的幽默传播载体。梗图通常包含图像与文字,可能被用于传播虚假信息或仇恨言论,因此对其进行细致研究至关重要。我们提出Memotion 3数据集,包含10,000张经过标注的梗图。与领域内现有主流数据集(包括Memotion早期版本)不同,Memotion 3引入印英混合语码梗图,而此前相关研究仅局限于英文梗图。本文描述了Memotion任务、数据收集及数据集构建方法,并提供了任务的基准模型。基准代码与数据集将在https://github.com/Shreyashm16/Memotion-3.0 公开。