Motivated by social network analysis and network-based recommendation systems, we study a semi-supervised community detection problem in which the objective is to estimate the community label of a new node using the network topology and partially observed community labels of existing nodes. The network is modeled using a degree-corrected stochastic block model, which allows for severe degree heterogeneity and potentially non-assortative communities. We propose an algorithm that computes a `structural similarity metric' between the new node and each of the $K$ communities by aggregating labeled and unlabeled data. The estimated label of the new node corresponds to the value of $k$ that maximizes this similarity metric. Our method is fast and numerically outperforms existing semi-supervised algorithms. Theoretically, we derive explicit bounds for the misclassification error and show the efficiency of our method by comparing it with an ideal classifier. Our findings highlight, to the best of our knowledge, the first semi-supervised community detection algorithm that offers theoretical guarantees.
翻译:受社交网络分析和基于网络的推荐系统启发,我们研究了一个半监督社区检测问题,其目标是通过网络拓扑结构以及已有节点的部分观测社区标签,来估计新节点的社区标签。网络采用度校正随机块模型进行建模,该模型允许存在显著的度异质性以及潜在的非同配社区。我们提出了一种算法,通过聚合有标签和无标签数据,计算新节点与每个社区(共K个)之间的“结构相似度指标”。新节点的估计标签对应于使该相似度指标最大化的k值。我们的方法速度快,且在数值上优于现有的半监督算法。在理论上,我们推导了误分类误差的显式界,并通过与理想分类器的比较展示了我们方法的有效性。据我们所知,我们的研究成果首次提出了一种提供理论保证的半监督社区检测算法。