It is often observed that stochastic gradient descent (SGD) and its variants implicitly select a solution with good generalization performance; such implicit bias is often characterized in terms of the sharpness of the minima. Kleinberg et al. (2018) connected this bias with the smoothing effect of SGD which eliminates sharp local minima by the convolution using the stochastic gradient noise. We follow this line of research and study the commonly-used averaged SGD algorithm, which has been empirically observed in Izmailov et al. (2018) to prefer a flat minimum and therefore achieves better generalization. We prove that in certain problem settings, averaged SGD can efficiently optimize the smoothed objective which avoids sharp local minima. In experiments, we verify our theory and show that parameter averaging with an appropriate step size indeed leads to significant improvement in the performance of SGD.
翻译:随机梯度下降(SGD)及其变体常被观察到能隐式地选择具有良好泛化性能的解;这种隐式偏差通常用极小点的锐度来刻画。Kleinberg等人(2018)将这种偏差与SGD的平滑效应联系起来,后者通过利用随机梯度噪声进行卷积来消除尖锐的局部极小点。我们沿着这一研究方向,探讨了常用的平均SGD算法。Izmailov等人(2018)通过实验观察到,该算法倾向于选择平坦的极小点,从而获得更好的泛化性能。我们证明,在某些问题设定下,平均SGD能够高效地优化平滑后的目标函数,从而避开尖锐的局部极小点。在实验中,我们验证了理论,并表明采用合适步长的参数平均确实能显著提升SGD的性能。