The discourse around conspiracy theories is currently thriving amidst the rampant misinformation in online environments. Research in this field has been focused on detecting conspiracy theories on social media, often relying on limited datasets. In this study, we present a novel methodology for constructing a Twitter dataset that encompasses accounts engaged in conspiracy-related activities throughout the year 2022. Our approach centers on data collection that is independent of specific conspiracy theories and information operations. Additionally, our dataset includes a control group comprising randomly selected users who can be fairly compared to the individuals involved in conspiracy activities. This comprehensive collection effort yielded a total of 15K accounts and 37M tweets extracted from their timelines. We conduct a comparative analysis of the two groups across three dimensions: topics, profiles, and behavioral characteristics. The results indicate that conspiracy and control users exhibit similarity in terms of their profile metadata characteristics. However, they diverge significantly in terms of behavior and activity, particularly regarding the discussed topics, the terminology used, and their stance on trending subjects. In addition, we find no significant disparity in the presence of bot users between the two groups. Finally, we develop a classifier to identify conspiracy users using features borrowed from bot, troll and linguistic literature. The results demonstrate a high accuracy level (with an F1 score of 0.94), enabling us to uncover the most discriminating features associated with conspiracy-related accounts.
翻译:当前,围绕阴谋论的话语在在线环境中虚假信息泛滥的背景下愈发活跃。该领域的研究集中于检测社交媒体上的阴谋论,但通常依赖有限的数据集。在本研究中,我们提出了一种新方法,用于构建一个涵盖2022年全年参与阴谋相关活动的推特账号数据集。我们的方法核心在于数据收集独立于特定阴谋论和信息操作。此外,我们的数据集包含一个由随机选取用户组成的对照组,这些用户可与涉及阴谋活动的个体进行公平比较。这一综合收集工作共产生了1.5万个账号及其时间线中提取的3700万条推文。我们从主题、个人资料和行为特征三个维度对两组进行比较分析。结果表明,阴谋用户与对照用户在个人资料元数据特征上表现出相似性,但在行为和活动方面存在显著差异,尤其是在讨论的主题、使用的术语以及对热门话题的立场上。此外,我们发现两组之间在机器人用户的存在上并无显著差异。最后,我们利用从机器人、网络喷子和语言学文献中借鉴的特征开发了一个分类器,用于识别阴谋用户。结果显示出高准确率(F1得分为0.94),使我们能够揭示与阴谋相关账号最具区分度的特征。