We show that for separable convex optimization, random stepsizes fully accelerate Gradient Descent. Specifically, using inverse stepsizes i.i.d. from the Arcsine distribution improves the convergence rate from $O(k)$ to $O(\sqrt{k})$, where $k$ is the condition number. No momentum or other algorithmic modifications are required. Our starting point is a remarkable "equalization property" of the Arcsine distribution: it yields an identical convergence rate for all quadratic functions. A key technical insight is that martingale arguments extend this phenomenon to all separable convex functions. We interpret this equalization as an extreme form of hedging: by using this random distribution over stepsizes, Gradient Descent converges at exactly the same rate for all functions in the function class.
翻译:我们证明,对于可分凸优化,随机步长能够完全加速梯度下降法。具体而言,使用服从反正弦分布的独立同分布逆步长,可将收敛率从$O(k)$提升至$O(\sqrt{k})$,其中$k$为条件数。该方法无需引入动量或其他算法改进。我们的出发点在于反正弦分布具有显著的“均衡特性”:它能使所有二次函数获得一致的收敛速率。关键技术洞察在于,鞅论证可将此现象推广至所有可分凸函数。我们将这种均衡性解释为对冲的极端形式:通过使用该随机步长分布,梯度下降法对函数类中所有函数的收敛速率完全一致。