Network clustering tackles the problem of identifying sets of nodes (communities) that have similar connection patterns. However, in many scenarios, nodes also have attributes that are correlated with the clustering structure. Thus, network information (edges) and node information (attributes) can be jointly leveraged to design high-performance clustering algorithms. Under a general model for the network and node attributes, this work establishes an information-theoretic criterion for the exact recovery of community labels and characterizes a phase transition determined by the Chernoff-Hellinger divergence of the model. The criterion shows how network and attribute information can be exchanged in order to have exact recovery (e.g., more reliable network information requires less reliable attribute information). This work also presents an iterative clustering algorithm that maximizes the joint likelihood, assuming that the probability distribution of network interactions and node attributes belong to exponential families. This covers a broad range of possible interactions (e.g., edges with weights) and attributes (e.g., non-Gaussian models), as well as sparse networks, while also exploring the connection between exponential families and Bregman divergences. Extensive numerical experiments using synthetic data indicate that the proposed algorithm outperforms classic algorithms that leverage only network or only attribute information as well as state-of-the-art algorithms that also leverage both sources of information. The contributions of this work provide insights into the fundamental limits and practical techniques for inferring community labels on node-attributed networks.
翻译:网络聚类解决识别具有相似连接模式的节点集(社区)的问题。然而在许多场景中,节点还具备与聚类结构相关的属性。因此,网络信息(边)与节点信息(属性)可被联合利用以设计高性能聚类算法。在针对网络与节点属性的通用模型下,本文建立了社区标签精确恢复的信息论准则,刻画了由模型的Chernoff-Hellinger散度决定的相变现象。该准则揭示了网络与属性信息如何相互替代以实现精确恢复(例如,更可靠的网络信息可降低对属性信息可靠性的要求)。本文还提出一种迭代聚类算法,通过最大化联合似然函数,假设网络交互与节点属性的概率分布属于指数族。这覆盖了广泛的交互形式(如带权边)与属性(如非高斯模型)及稀疏网络,同时探索了指数族与Bregman散度之间的关联。基于合成数据的大量数值实验表明,所提算法优于仅利用网络或仅利用属性信息的经典算法,以及同时利用两类信息的最先进算法。本文的贡献为理解节点属性网络中社区标签推断的理论极限与实用技术提供了洞见。