We propose novel optimal and parameter-free algorithms for computing an approximate solution for smooth optimization with small (projected) gradient norm. Specifically, for computing an approximate solution such that the norm of the (projected) gradient is not greater than $\varepsilon$, we have the following results for the cases of convex, strongly convex, and nonconvex problems: a) for the convex case, the total number of gradient evaluations is bounded by $O(1)\sqrt{L\|x_0 - x^*\|\varepsilon}$, where $L$ is the Lipschitz constant of the gradient function and $x^*$ is any optimal solution; b) for the strongly convex case, the total number of gradient evaluations is bounded by $O(1)\sqrt{L/\mu}\log(\|\nabla f(x_0)\|)$, where $\mu$ is the strong convexity constant; c) for the nonconvex case, the total number of gradient evaluations is bounded by $O(1)\sqrt{Ll}(f(x_0) - f(x^*))/\varepsilon^2$, where $l$ is the lower curvature constant. Our complexity results match the lower complexity bounds of all three cases of problems. Our analysis can be applied to both unconstrained problems and problems with constrained feasible sets; we demonstrate our strategy for analyzing the complexity of computing solutions with small projected gradient norm in the convex case. For all the convex, strongly convex, and nonconvex cases, we also propose parameter-free algorithms that does not require the knowledge of any problem parameter. To the best of our knowledge, our paper is the first one that achieves the $O(1)\sqrt{L\|x_0 - x^*\|/\varepsilon}$ complexity for convex problems with constraint feasible sets, the $O(1)\sqrt{Ll}(f(x_0) - f(x^*))/\varepsilon$ complexity for nonconvex problems, and optimal complexities for convex, strongly convex, and nonconvex problems through parameter-free algorithms.
翻译:本文针对光滑优化问题,提出了一种新颖的最优且无需参数的算法,用于计算具有较小(投影)梯度范数的近似解。具体而言,在计算满足(投影)梯度范数不超过$\varepsilon$的近似解时,对于凸、强凸和非凸问题,我们获得以下结果:a)对于凸情形,梯度评估总次数受限于$O(1)\sqrt{L\|x_0 - x^*\|\varepsilon}$,其中$L$为梯度函数的Lipschitz常数,$x^*$为任意最优解;b)对于强凸情形,梯度评估总次数受限于$O(1)\sqrt{L/\mu}\log(\|\nabla f(x_0)\|)$,其中$\mu$为强凸常数;c)对于非凸情形,梯度评估总次数受限于$O(1)\sqrt{Ll}(f(x_0) - f(x^*))/\varepsilon^2$,其中$l$为下曲率常数。我们的复杂度结果与三类问题的下界复杂度相匹配。本文分析可同时适用于无约束问题和带约束可行集问题;我们展示了在凸情形下分析具有小投影梯度范数解的计算复杂度的策略。针对凸、强凸和非凸所有情形,我们还提出了无需任何问题参数知识的无参数算法。据我们所知,本文首次通过无参数算法实现了:带约束可行集凸问题的$O(1)\sqrt{L\|x_0 - x^*\|/\varepsilon}$复杂度、非凸问题的$O(1)\sqrt{Ll}(f(x_0) - f(x^*))/\varepsilon$复杂度,以及凸、强凸和非凸问题的最优复杂度。