Weak recovery, hypothesis testing, and mutual information in stochastic block models and planted factor graphs

The stochastic block model is a canonical model of communities in random graphs. It was introduced in the social sciences and statistics as a model of communities, and in theoretical computer science as an average case model for graph partitioning problems under the name of the ``planted partition model.'' Given a sparse stochastic block model, the two standard inference tasks are: (i) Weak recovery: can we estimate the communities with non trivial overlap with the true communities? (ii) Detection/Hypothesis testing: can we distinguish if the sample was drawn from the block model or from a random graph with no community structure with probability tending to $1$ as the graph size tends to infinity? In this work, we show that for sparse stochastic block models, the two inference tasks are equivalent except at a critical point. That is, weak recovery is information theoretically possible if and only if detection is possible. We thus find a strong connection between these two notions of inference for the model. We further prove that when detection is impossible, an explicit hypothesis test based on low degree polynomials in the adjacency matrix of the observed graph achieves the optimal statistical power. This low degree test is efficient as opposed to the likelihood ratio test, which is not known to be efficient. Moreover, we prove that the asymptotic mutual information between the observed network and the community structure exhibits a phase transition at the weak recovery threshold. Our results are proven in much broader settings including the hypergraph stochastic block models and general planted factor graphs. In these settings we prove that the impossibility of weak recovery implies contiguity and provide a condition which guarantees the equivalence of weak recovery and detection.

翻译：随机块模型是随机图中社区结构的经典模型。该模型最初在社会科学与统计学中被提出，用于描述社区结构；在理论计算机科学中，它作为图划分问题的平均情况模型，被称为"植入划分模型"。给定一个稀疏随机块模型，两个标准的推断任务是：(i) 弱恢复：能否以非平凡的重合度估计真实社区？(ii) 检测/假设检验：当图规模趋于无穷时，能否以概率趋于$1$区分样本是来自块模型还是来自无社区结构的随机图？本研究表明，对于稀疏随机块模型，除临界点外，这两个推断任务是等价的。也就是说，弱恢复在信息论意义上可能当且仅当检测可能。由此我们揭示了该模型两种推断概念之间的深刻联系。我们进一步证明，当检测不可能时，基于观测图邻接矩阵低阶多项式的显式假设检验能达到最优统计功效。与似然比检验（其高效性尚未知）不同，这种低阶检验是高效的。此外，我们证明了观测网络与社区结构之间的渐近互信息在弱恢复阈值处呈现相变。我们的结果在更广泛的设定下成立，包括超图随机块模型和一般植入因子图。在这些设定中，我们证明了弱恢复的不可能性意味着连续性，并给出了保证弱恢复与检测等价的条件。

相关内容

MoDELS

关注 46

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

《生成式模型: 变分自编码器与扩散模型》，75页ppt，Google DeepMind科学家Ruiqi Gao

专知会员服务

66+阅读 · 2023年6月10日

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

分布外泛化(Out-Of-Distribution Generalization) 综述论文，22页pdf240篇文献

专知会员服务

64+阅读 · 2021年9月2日