We characterize the learning dynamics of stochastic gradient descent (SGD) when continuous symmetry exists in the loss function, where the divergence between SGD and gradient descent is dramatic. We show that depending on how the symmetry affects the learning dynamics, we can divide a family of symmetry into two classes. For one class of symmetry, SGD naturally converges to solutions that have a balanced and aligned gradient noise. For the other class of symmetry, SGD will almost always diverge. Then, we show that our result remains applicable and can help us understand the training dynamics even when the symmetry is not present in the loss function. Our main result is universal in the sense that it only depends on the existence of the symmetry and is independent of the details of the loss function. We demonstrate that the proposed theory offers an explanation of progressive sharpening and flattening and can be applied to common practical problems such as representation normalization, matrix factorization, and the use of warmup.
翻译:我们刻画了损失函数存在连续对称性时随机梯度下降(SGD)的学习动力学特征,此时SGD与梯度下降的差异极为显著。研究表明,根据对称性对学习动力学的影响方式,可将对称性家族划分为两类:对于第一类对称性,SGD自然收敛至具有平衡且对齐梯度噪声的解;对于第二类对称性,SGD几乎必然发散。进一步,我们发现即便损失函数中不存在对称性,该结论仍然成立,并可助力理解训练动力学过程。本文主要结论具有普适性,仅依赖于对称性的存在,而与损失函数的具体形式无关。实验证明,该理论可解释渐进锐化与平坦化现象,并适用于表示归一化、矩阵分解及预热策略等常见实际场景。