The COVID-19 pandemic led to an infodemic where an overwhelming amount of COVID-19 related content was being disseminated at high velocity through social media. This made it challenging for citizens to differentiate between accurate and inaccurate information about COVID-19. This motivated us to carry out a comparative study of the characteristics of COVID-19 misinformation versus those of accurate COVID-19 information through a large-scale computational analysis of over 242 million tweets. The study makes comparisons alongside four key aspects: 1) the distribution of topics, 2) the live status of tweets, 3) language analysis and 4) the spreading power over time. An added contribution of this study is the creation of a COVID-19 misinformation classification dataset. Finally, we demonstrate that this new dataset helps improve misinformation classification by more than 9\% based on average F1 measure.
翻译:COVID-19疫情引发了信息疫情,大量与COVID-19相关的内容通过社交媒体高速传播,使公众难以区分关于COVID-19的准确与不准确信息。这促使我们对超过2.42亿条推文开展大规模计算分析,以比较COVID-19虚假信息与准确信息的特征。本研究从四个关键维度进行比较:1)主题分布,2)推文活跃状态,3)语言分析,4)随时间推移的传播力。本研究的另一贡献是构建了COVID-19虚假信息分类数据集。最后,我们证明该新数据集可使基于平均F1值的虚假信息分类性能提升超过9%。