Higher-order structures of networks, namely, small subgraphs of networks (also called network motifs), are widely known to be crucial and essential to the organization of networks. There has been a few work studying the community detection problem -- a fundamental problem in network analysis, at the level of motifs. In particular, higher-order spectral clustering has been developed, where the notion of motif adjacency matrix is introduced as the input of the algorithm. However, it remains largely unknown that how higher-order spectral clustering works and when it performs better than its edge-based counterpart. To elucidate these problems, we investigate higher-order spectral clustering from a statistical perspective. In particular, we theoretically study the clustering performance of higher-order spectral clustering under a weighted stochastic block model and compare the resulting bounds with the corresponding results of edge-based spectral clustering. It turns out that when the network is dense with weak signal of weights, higher-order spectral clustering can really lead to the performance gain in clustering. We also use simulations and real data experiments to support the findings.
翻译:网络的高阶结构,即网络中的小规模子图(也称为网络模体),已被广泛认为对网络组织具有关键性和必要性。已有少量工作研究了模体层面的社区检测问题——网络分析中的一个基础问题。特别是,高阶谱聚类方法已被提出,其中引入了模体邻接矩阵作为算法的输入。然而,高阶谱聚类如何工作以及何时比基于边的谱聚类更优仍基本未知。为阐明这些问题,我们从统计视角研究高阶谱聚类。具体而言,我们在加权随机块模型下理论分析高阶谱聚类的聚类性能,并将所得界与基于边的谱聚类的对应结果进行比较。研究发现,当网络密集且权重信号较弱时,高阶谱聚类确实能带来聚类性能的提升。我们还通过仿真和真实数据实验支持这些发现。