The use of propaganda has spiked on mainstream and social media, aiming to manipulate or mislead users. While efforts to automatically detect propaganda techniques in textual, visual, or multimodal content have increased, most of them primarily focus on English content. The majority of the recent initiatives targeting medium to low-resource languages produced relatively small annotated datasets, with a skewed distribution, posing challenges for the development of sophisticated propaganda detection models. To address this challenge, we carefully develop the largest propaganda dataset to date, ArPro, comprised of 8K paragraphs from newspaper articles, labeled at the text span level following a taxonomy of 23 propagandistic techniques. Furthermore, our work offers the first attempt to understand the performance of large language models (LLMs), using GPT-4, for fine-grained propaganda detection from text. Results showed that GPT-4's performance degrades as the task moves from simply classifying a paragraph as propagandistic or not, to the fine-grained task of detecting propaganda techniques and their manifestation in text. Compared to models fine-tuned on the dataset for propaganda detection at different classification granularities, GPT-4 is still far behind. Finally, we evaluate GPT-4 on a dataset consisting of six other languages for span detection, and results suggest that the model struggles with the task across languages. Our dataset and resources will be released to the community.
翻译:宣传手段在主流媒体和社交媒体上的使用激增,旨在操纵或误导用户。尽管自动检测文本、视觉或多模态内容中宣传技术的努力有所增加,但大多数研究主要集中在英语内容上。近期针对中低资源语言的项目大多产生了规模较小且分布不均的标注数据集,给开发复杂的宣传检测模型带来了挑战。为应对这一挑战,我们精心构建了迄今为止最大的宣传数据集ArPro,包含来自报纸文章的8000个段落,按照23种宣传技术的分类体系在文本片段层面进行标注。此外,本研究首次尝试理解大型语言模型(以GPT-4为例)在细粒度文本宣传检测中的表现。结果表明,当任务从简单判断段落是否具有宣传性,转向检测具体宣传技术及其在文本中体现的细粒度任务时,GPT-4的性能显著下降。与在不同分类粒度上针对宣传检测进行微调的模型相比,GPT-4仍存在较大差距。最后,我们在包含六种其他语言的文本片检测数据集上评估GPT-4,结果表明该模型在多语言场景下均难以胜任该任务。我们的数据集和资源将向社区开放。