We consider alternating gradient descent (AGD) with fixed step size applied to the asymmetric matrix factorization objective. We show that, for a rank-$r$ matrix $\mathbf{A} \in \mathbb{R}^{m \times n}$, $T = C (\frac{\sigma_1(\mathbf{A})}{\sigma_r(\mathbf{A})})^2 \log(1/\epsilon)$ iterations of alternating gradient descent suffice to reach an $\epsilon$-optimal factorization $\| \mathbf{A} - \mathbf{X} \mathbf{Y}^{T} \|^2 \leq \epsilon \| \mathbf{A}\|^2$ with high probability starting from an atypical random initialization. The factors have rank $d \geq r$ so that $\mathbf{X}_{T}\in\mathbb{R}^{m \times d}$ and $\mathbf{Y}_{T} \in\mathbb{R}^{n \times d}$, and mild overparameterization suffices for the constant $C$ in the iteration complexity $T$ to be an absolute constant. Experiments suggest that our proposed initialization is not merely of theoretical benefit, but rather significantly improves the convergence rate of gradient descent in practice. Our proof is conceptually simple: a uniform Polyak-\L{}ojasiewicz (PL) inequality and uniform Lipschitz smoothness constant are guaranteed for a sufficient number of iterations, starting from our random initialization. Our proof method should be useful for extending and simplifying convergence analyses for a broader class of nonconvex low-rank factorization problems.
翻译:我们研究了固定步长的交替梯度下降(AGD)在非对称矩阵分解目标函数上的应用。我们证明,对于一个秩为$r$的矩阵$\mathbf{A} \in \mathbb{R}^{m \times n}$,从非典型随机初始化出发,交替梯度下降算法经过$T = C (\frac{\sigma_1(\mathbf{A})}{\sigma_r(\mathbf{A})})^2 \log(1/\epsilon)$次迭代,即可高概率达到$\epsilon$-最优分解$\| \mathbf{A} - \mathbf{X} \mathbf{Y}^{T} \|^2 \leq \epsilon \| \mathbf{A}\|^2$。分解因子具有秩$d \geq r$,从而$\mathbf{X}_{T}\in\mathbb{R}^{m \times d}$且$\mathbf{Y}_{T} \in\mathbb{R}^{n \times d}$,且适度的过参数化足以使得迭代复杂度$T$中的常数$C$成为一个绝对常数。实验表明,我们所提出的初始化方法不仅具有理论价值,而且在实际中显著提升了梯度下降的收敛速度。我们的证明概念简洁:从随机初始化开始,在足够多的迭代次数内,可以保证一致的Polyak-Łojasiewicz(PL)不等式和均匀Lipschitz平滑常数。我们的证明方法应当有助于推广并简化更广泛一类非凸低秩分解问题的收敛性分析。