We consider the problem of learning a sparse graph under the Laplacian constrained Gaussian graphical models. This problem can be formulated as a penalized maximum likelihood estimation of the Laplacian constrained precision matrix. Like in the classical graphical lasso problem, recent works made use of the $\ell_1$-norm regularization with the goal of promoting sparsity in Laplacian constrained precision matrix estimation. However, we find that the widely used $\ell_1$-norm is not effective in imposing a sparse solution in this problem. Through empirical evidence, we observe that the number of nonzero graph weights grows with the increase of the regularization parameter. From a theoretical perspective, we prove that a large regularization parameter will surprisingly lead to a complete graph, i.e., every pair of vertices is connected by an edge. To address this issue, we introduce the nonconvex sparsity penalty, and propose a new estimator by solving a sequence of weighted $\ell_1$-norm penalized sub-problems. We establish the non-asymptotic optimization performance guarantees on both optimization error and statistical error, and prove that the proposed estimator can recover the edges correctly with a high probability. To solve each sub-problem, we develop a projected gradient descent algorithm which enjoys a linear convergence rate. Finally, an extension to learn disconnected graphs is proposed by imposing additional rank constraint. We propose a numerical algorithm based on based on the alternating direction method of multipliers, and establish its theoretical sequence convergence. Numerical experiments involving synthetic and real-world data sets demonstrate the effectiveness of the proposed method.
翻译:摘要:我们研究在拉普拉斯约束高斯图模型下学习稀疏图的问题。该问题可表述为对拉普拉斯约束精度矩阵的惩罚最大似然估计。与经典图形套索问题类似,近期研究利用$\ell_1$-范数正则化以促进拉普拉斯约束精度矩阵估计的稀疏性。然而,我们发现广泛使用的$\ell_1$-范数在此问题中难以有效施加稀疏解。通过实证证据,我们观察到非零图权重的数量随正则化参数增大而增加。从理论角度,我们证明较大的正则化参数将出乎意料地导致完全图,即每对顶点均由边连接。为解决此问题,我们引入非凸稀疏惩罚,并通过求解一系列加权$\ell_1$-范数惩罚子问题提出新估计量。我们建立了优化误差与统计误差的非渐近优化性能保证,并证明所提估计量能以高概率正确恢复边。为求解每个子问题,我们开发了具有线性收敛速率的投影梯度下降算法。最后,通过施加额外秩约束,提出学习不连通图的扩展方法。我们设计了基于交替方向乘子法的数值算法,并建立其理论序列收敛性。涉及合成与真实数据集的数值实验证明了所提方法的有效性。