Recently, several studies consider the stochastic optimization problem but in a heavy-tailed noise regime, i.e., the difference between the stochastic gradient and the true gradient is assumed to have a finite $p$-th moment (say being upper bounded by $\sigma^{p}$ for some $\sigma\geq0$) where $p\in(1,2]$, which not only generalizes the traditional finite variance assumption ($p=2$) but also has been observed in practice for several different tasks. Under this challenging assumption, lots of new progress has been made for either convex or nonconvex problems, however, most of which only consider smooth objectives. In contrast, people have not fully explored and well understood this problem when functions are nonsmooth. This paper aims to fill this crucial gap by providing a comprehensive analysis of stochastic nonsmooth convex optimization with heavy-tailed noises. We revisit a simple clipping-based algorithm, whereas, which is only proved to converge in expectation but under the additional strong convexity assumption. Under appropriate choices of parameters, for both convex and strongly convex functions, we not only establish the first high-probability rates but also give refined in-expectation bounds compared with existing works. Remarkably, all of our results are optimal (or nearly optimal up to logarithmic factors) with respect to the time horizon $T$ even when $T$ is unknown in advance. Additionally, we show how to make the algorithm parameter-free with respect to $\sigma$, in other words, the algorithm can still guarantee convergence without any prior knowledge of $\sigma$. Furthermore, an initial distance adaptive convergence rate is provided if $\sigma$ is assumed to be known.
翻译:摘要:近年来,多项研究考虑了随机优化问题,但处于重尾噪声框架下(即随机梯度与真实梯度之差假定具有有限的$p$阶矩,例如对于某$\sigma\geq0$存在上界$\sigma^{p}$,其中$p\in(1,2]$),这不仅推广了传统有限方差假设($p=2$),而且已在多个不同任务的实际观测中得到验证。在此困难假设下,针对凸或非凸问题已取得诸多新进展,然而多数工作仅考虑光滑目标。相比之下,当函数非光滑时,人们对这一问题的探索尚不充分且理解不足。本文旨在通过系统分析带重尾噪声的随机非光滑凸优化来填补这一关键空白。我们重新审视一种基于裁剪的简单算法,此前该算法仅在额外强凸假设下证明具有期望收敛性。在参数适当选取下,对于凸函数和强凸函数,我们不仅建立了首个高概率收敛率,还给出了相较现有工作更精细的期望界。值得注意的是,即便时间跨度$T$预先未知,我们所有结果关于$T$均为最优(或对数因子下近最优)。此外,我们展示了如何使算法在$\sigma$方面实现无参数化——即无需$\sigma$任何先验知识,算法仍能保证收敛。进一步,若假定$\sigma$已知,则能推导出初始距离自适应收敛率。