We consider using gradient descent to minimize the nonconvex function $f(X)=\phi(XX^{T})$ over an $n\times r$ factor matrix $X$, in which $\phi$ is an underlying smooth convex cost function defined over $n\times n$ matrices. While only a second-order stationary point $X$ can be provably found in reasonable time, if $X$ is additionally rank deficient, then its rank deficiency certifies it as being globally optimal. This way of certifying global optimality necessarily requires the search rank $r$ of the current iterate $X$ to be overparameterized with respect to the rank $r^{\star}$ of the global minimizer $X^{\star}$. Unfortunately, overparameterization significantly slows down the convergence of gradient descent, from a linear rate with $r=r^{\star}$ to a sublinear rate when $r>r^{\star}$, even when $\phi$ is strongly convex. In this paper, we propose an inexpensive preconditioner that restores the convergence rate of gradient descent back to linear in the overparameterized case, while also making it agnostic to possible ill-conditioning in the global minimizer $X^{\star}$.
翻译:本文考虑使用梯度下降法最小化定义在$n\times r$因子矩阵$X$上的非凸函数$f(X)=\phi(XX^{T})$,其中$\phi$是定义在$n\times n$矩阵上的底层光滑凸代价函数。虽然理论上可在合理时间内找到二阶稳定点$X$,但若$X$额外满足秩亏缺条件,则其秩亏缺性可保证该点为全局最优解。这种全局最优性的认证方式要求当前迭代点$X$的搜索秩$r$相对于全局最优解$X^{\star}$的秩$r^{\star}$具有过参数化特征。然而,过参数化会显著降低梯度下降法的收敛速度:当$\phi$为强凸函数时,收敛速率从$r=r^{\star}$情况下的线性收敛退化为$r>r^{\star}$情况下的次线性收敛。本文提出了一种廉价预处理方法,可在过参数化情形下将梯度下降法的收敛速度恢复为线性收敛,同时使该方法对全局最优解$X^{\star}$可能存在的病态条件保持不敏感。