Symmetries exist abundantly in the loss function of neural networks. We characterize the learning dynamics of stochastic gradient descent (SGD) when exponential symmetries, a broad subclass of continuous symmetries, exist in the loss function. We establish that when gradient noises do not balance, SGD has the tendency to move the model parameters toward a point where noises from different directions are balanced. Here, a special type of fixed point in the constant directions of the loss function emerges as a candidate for solutions for SGD. As the main theoretical result, we prove that every parameter $\theta$ connects without loss function barrier to a unique noise-balanced fixed point $\theta^*$. The theory implies that the balancing of gradient noise can serve as a novel alternative mechanism for relevant phenomena such as progressive sharpening and flattening and can be applied to understand common practical problems such as representation normalization, matrix factorization, warmup, and formation of latent representations.
翻译:神经网络损失函数中广泛存在对称性。当损失函数中存在指数对称性(连续对称性的一个广泛子类)时,我们刻画了随机梯度下降(SGD)的学习动态。我们证明,当梯度噪声不平衡时,SGD倾向于将模型参数移向一个来自不同方向的噪声达到平衡的点。在此,损失函数常数方向上的一种特殊类型不动点作为SGD解的候选者出现。作为主要理论结果,我们证明了每个参数$\theta$均可在无损失函数障碍的情况下连接到一个唯一的噪声平衡不动点$\theta^*$。该理论表明,梯度噪声的平衡可以作为一种新颖的替代机制来解释渐进锐化与平坦化等相关现象,并可应用于理解表示归一化、矩阵分解、预热训练及潜在表示形成等常见实际问题。