In recent years, there has been a heightened consensus within academia and in the public discourse that Social Media Platforms (SMPs), amplify the spread of hateful and negative sentiment content. Researchers have identified how hateful content, political propaganda, and targeted messaging contributed to real-world harms including insurrections against democratically elected governments, genocide, and breakdown of social cohesion due to heightened negative discourse towards certain communities in parts of the world. To counter these issues, SMPs have created semi-automated systems that can help identify toxic speech. In this paper we analyse the statistical distribution of hateful and negative sentiment contents within a representative Facebook dataset (n= 604,703) scrapped through 648 public Facebook pages which identify themselves as proponents (and followers) of far-right Hindutva actors. These pages were identified manually using keyword searches on Facebook and on CrowdTangleand classified as far-right Hindutva pages based on page names, page descriptions, and discourses shared on these pages. We employ state-of-the-art, open-source XLM-T multilingual transformer-based language models to perform sentiment and hate speech analysis of the textual contents shared on these pages over a period of 5.5 years. The result shows the statistical distributions of the predicted sentiment and the hate speech labels; top actors, and top page categories. We further discuss the benchmark performances and limitations of these pre-trained language models.
翻译:近年来,学术界与公共领域日益形成共识:社交媒体平台(SMPs)加剧了仇恨性及负面情感内容的传播。研究者已证实,仇恨内容、政治宣传与定向信息如何对全球部分地区造成现实危害,包括对民选政府的暴乱、种族灭绝,以及因针对特定社群的负面言论激增而引发的社会凝聚力瓦解。为应对这些问题,社交媒体平台开发了半自动化系统以协助识别毒性言论。本文分析了代表性Facebook数据集(n=604,703)中仇恨与负面情感内容的统计分布,该数据通过648个公开Facebook页面采集,这些页面自称为极右翼印度教特性(Hindutva)行动者的支持者(及关注者)。我们通过Facebook和CrowdTangle的关键词搜索人工识别这些页面,并根据页面名称、描述及分享内容将其归类为极右翼Hindutva页面。采用当前最先进的开源XLM-T多语言Transformer语言模型,对上述页面在5.5年期间分享的文本内容进行情感与仇恨言论分析。结果展示了预测情感与仇恨言论标签的统计分布、主要行动者及页面类别。我们进一步讨论了这些预训练语言模型的基准性能及其局限性。