Community detection algorithms try to extract a mesoscale structure from the available network data, generally avoiding any explicit assumption regarding the quantity and quality of information conveyed by specific sets of edges. In this paper, we show that the core of ideological/discursive communities on X/Twitter can be effectively identified by uncovering the most informative interactions in an authors-audience bipartite network through a maximum-entropy null model. The analysis is performed considering three X/Twitter datasets related to the main political events of 2022 in Italy, using as benchmarks four state-of-the-art algorithms - three descriptive, one inferential -, and manually annotating nearly 300 verified users based on their political affiliation. In terms of information content, the communities obtained with the entropy-based algorithm are comparable to those obtained with some of the benchmarks. However, such a methodology on the authors-audience bipartite network: uses just a small sample of the available data to identify the central users of each community; returns a neater partition of the user set in just a few, easy to interpret, communities; clusters well-known political figures in a way that better matches the political alliances when compared with the benchmarks. Our results provide an important insight into online debates, highlighting that online interaction networks are mostly shaped by the activity of a small set of users who enjoy public visibility even outside social media.
翻译:社群检测算法试图从可用的网络数据中提取中尺度结构,通常避免对特定边集所传达信息的数量和质量做出明确假设。本文表明,通过最大熵零模型揭示作者-受众二分网络中最具信息量的交互,可以有效识别X/Twitter上意识形态/讨论社群的核心。分析基于三个与2022年意大利主要政治事件相关的X/Twitter数据集展开,以四种最先进的算法(三种描述性算法、一种推断性算法)为基准,并对近300名根据其政治归属手动标注的经过验证用户进行研究。在信息内容方面,基于熵的算法所得到的社群与部分基准算法所得社群具有可比性。然而,这种基于作者-受众二分网络的方法具有以下特点:仅使用可用数据中的一小部分样本即可识别每个社群的核心用户;将用户集划分为更清晰、数量较少且易于解释的社群;在聚类知名政治人物时,相较于基准算法更符合政治联盟的实际情况。我们的结果为在线辩论提供了重要洞见,凸显出在线交互网络主要由一小群甚至在社交媒体之外也享有公众知名度的用户的活动所塑造。