Whereas Laplacian and modularity based spectral clustering is apt to dense graphs, recent results show that for sparse ones, the non-backtracking spectrum is the best candidate to find assortative clusters of nodes. Here belief propagation in the sparse stochastic block model is derived with arbitrary given model parameters that results in a non-linear system of equations; with linear approximation, the spectrum of the non-backtracking matrix is able to specify the number $k$ of clusters. Then the model parameters themselves can be estimated by the EM algorithm. Bond percolation in the assortative model is considered in the following two senses: the within- and between-cluster edge probabilities decrease with the number of nodes and edges coming into existence in this way are retained with probability $\beta$. As a consequence, the optimal $k$ is the number of the structural real eigenvalues (greater than $\sqrt{c}$, where $c$ is the average degree) of the non-backtracking matrix of the graph. Assuming, these eigenvalues $\mu_1 >\dots > \mu_k$ are distinct, the multiple phase transitions obtained for $\beta$ are $\beta_i =\frac{c}{\mu_i^2}$; further, at $\beta_i$ the number of detectable clusters is $i$, for $i=1,\dots ,k$. Inflation-deflation techniques are also discussed to classify the nodes themselves, which can be the base of the sparse spectral clustering.
翻译:尽管基于拉普拉斯和模块度的谱聚类适用于稠密图,但最新结果表明,对于稀疏图而言,非回溯谱是发现节点同配聚类结构的最佳候选方案。本文推导了具有任意给定模型参数的稀疏随机分块模型中的置信传播算法,得到一组非线性方程组;通过线性近似,非回溯矩阵的谱能够确定聚类数目$k$。随后模型参数本身可通过EM算法估计。在同配模型中的键渗流以下述两种意义被考虑:簇内和簇间边概率随节点数减少,且以这种方式出现的边以概率$\beta$被保留。由此,最优$k$即为图非回溯矩阵中大于$\sqrt{c}$(其中$c$为平均度)的结构实特征值的个数。假设这些特征值$\mu_1 >\dots > \mu_k$互异,则针对$\beta$获得的多个相变点为$\beta_i =\frac{c}{\mu_i^2}$;进一步,在$\beta_i$处可检测的聚类数目为$i$,其中$i=1,\dots ,k$。本文还讨论了用于节点分类的膨胀-压缩技术,这可以作为稀疏谱聚类的基础。