We study the bias of Stochastic Gradient Descent (SGD) to learn low-rank weight matrices when training deep neural networks. Our results show that training neural networks with mini-batch SGD and weight decay causes a bias towards rank minimization over the weight matrices. Specifically, we show, both theoretically and empirically, that this bias is more pronounced when using smaller batch sizes, higher learning rates, or increased weight decay. Additionally, we predict and observe empirically that weight decay is necessary to achieve this bias. Unlike previous literature, our analysis does not rely on assumptions about the data, convergence, or optimality of the weight matrices and applies to a wide range of neural network architectures of any width or depth. Finally, we empirically investigate the connection between this bias and generalization, finding that it has a marginal effect on generalization.
翻译:我们研究了随机梯度下降法(SGD)在训练深度神经网络时学习低秩权重矩阵的偏差。结果表明,使用小批量SGD和权重衰减训练神经网络会促使权重矩阵向秩最小化方向产生偏差。具体而言,我们从理论和实验两方面证明,这种偏差在使用更小批量尺寸、更高学习率或更强权重衰减时更为显著。此外,我们通过理论预测与实验观察证实,权重衰减是实现该偏差的必要条件。与以往文献不同,我们的分析不依赖关于数据、收敛性或权重矩阵最优性的假设,且适用于任意宽度或深度的广泛神经网络架构。最后,我们通过实证研究了该偏差与泛化性能之间的关联,发现其对泛化能力的影响较为微弱。