The ongoing Russo-Ukrainian conflict has been a subject of intense media coverage worldwide. Understanding the global narrative surrounding this topic is crucial for researchers that aim to gain insights into its multifaceted dimensions. In this paper, we present a novel dataset that focuses on this topic by collecting and processing tweets posted by news or media companies on social media across the globe. We collected tweets from February 2022 to May 2023 to acquire approximately 1.5 million tweets in 60 different languages. Each tweet in the dataset is accompanied by processed tags, allowing for the identification of entities, stances, concepts, and sentiments expressed. The availability of the dataset serves as a valuable resource for researchers aiming to investigate the global narrative surrounding the ongoing conflict from various aspects such as who are the prominent entities involved, what stances are taken, where do these stances originate, and how are the different concepts related to the event portrayed.
翻译:持续的俄乌冲突一直是全球媒体密集报道的主题。理解围绕该主题的全球叙事,对于旨在洞察其多维层面的研究人员至关重要。本文介绍了一个聚焦于此主题的新型数据集,我们收集并处理了全球新闻或媒体公司在社交媒体上发布的推文。我们收集了2022年2月至2023年5月间的推文,获取了约150万条涉及60种不同语言的推文。数据集中的每条推文都附有处理后的标签,便于识别所表达的实体、立场、概念和情感。该数据集的可用性为研究人员提供了宝贵资源,使其能够从多个方面调查围绕这场持续冲突的全球叙事,例如涉及的主要实体有哪些、采取了何种立场、这些立场源自何处,以及事件相关的不同概念是如何被描绘的。