In this paper, we study the bias of Stochastic Gradient Descent (SGD) to learn low-rank weight matrices when training deep ReLU neural networks. Our results show that training neural networks with mini-batch SGD and weight decay causes a bias towards rank minimization over the weight matrices. Specifically, we show, both theoretically and empirically, that this bias is more pronounced when using smaller batch sizes, higher learning rates, or increased weight decay. Additionally, we predict and observe empirically that weight decay is necessary to achieve this bias. Finally, we empirically investigate the connection between this bias and generalization, finding that it has a marginal effect on generalization. Our analysis is based on a minimal set of assumptions and applies to neural networks of any width or depth, including those with residual connections and convolutional layers.
翻译:本文研究了在训练深层ReLU神经网络时,随机梯度下降(SGD)学习低秩权重矩阵的偏置。我们的结果表明,使用小批量SGD和权重衰减训练神经网络会导致权重矩阵向秩最小化的偏置。具体而言,我们通过理论和实验证明,当使用更小的批量大小、更高的学习率或更强的权重衰减时,这种偏置更为显著。此外,我们预测并通过实验观察到,权重衰减是实现这一偏置的必要条件。最后,我们通过实证研究了该偏置与泛化性能之间的关系,发现其对泛化影响有限。我们的分析基于最少的假设,适用于任意宽度或深度的神经网络,包括具有残差连接和卷积层的网络。