This paper investigates the language of propaganda and its stylistic features. It presents the PPN dataset, standing for Propagandist Pseudo-News, a multisource, multilingual, multimodal dataset composed of news articles extracted from websites identified as propaganda sources by expert agencies. A limited sample from this set was randomly mixed with papers from the regular French press, and their URL masked, to conduct an annotation-experiment by humans, using 11 distinct labels. The results show that human annotators were able to reliably discriminate between the two types of press across each of the labels. We propose different NLP techniques to identify the cues used by the annotators, and to compare them with machine classification. They include the analyzer VAGO to measure discourse vagueness and subjectivity, a TF-IDF to serve as a baseline, and four different classifiers: two RoBERTa-based models, CATS using syntax, and one XGBoost combining syntactic and semantic features.
翻译:本文研究了宣传语言及其风格特征。我们提出了PPN数据集,全称为“宣传类伪新闻”(Propagandist Pseudo-News),这是一个多源、多语言、多模态的数据集,包含从专家机构认定的宣传源网站中提取的新闻文章。从该数据集中随机抽取有限样本,与法国正规新闻报刊的文章混合,并隐藏其URL链接,采用11种不同标签进行人工标注实验。结果表明,人类标注者能够可靠地区分两种新闻类型,且各标签均表现一致。我们提出多种自然语言处理技术,用于识别标注者所依赖的语言线索,并将其与机器分类结果进行对比。这些技术包括:用于测量话语模糊性和主观性的VAGO分析器、作为基准的TF-IDF模型,以及四种分类器——两种基于RoBERTa的模型、基于句法的CATS模型,以及融合句法和语义特征的XGBoost模型。