We study benign overfitting in two-layer ReLU networks trained using gradient descent and hinge loss on noisy data for binary classification. In particular, we consider linearly separable data for which a relatively small proportion of labels are corrupted or flipped. We identify conditions on the margin of the clean data that give rise to three distinct training outcomes: benign overfitting, in which zero loss is achieved and with high probability test data is classified correctly; overfitting, in which zero loss is achieved but test data is misclassified with probability lower bounded by a constant; and non-overfitting, in which clean points, but not corrupt points, achieve zero loss and again with high probability test data is classified correctly. Our analysis provides a fine-grained description of the dynamics of neurons throughout training and reveals two distinct phases: in the first phase clean points achieve close to zero loss, in the second phase clean points oscillate on the boundary of zero loss while corrupt points either converge towards zero loss or are eventually zeroed by the network. We prove these results using a combinatorial approach that involves bounding the number of clean versus corrupt updates across these phases of training.
翻译:我们研究了在二分类含噪数据上,使用梯度下降和铰链损失训练的两层ReLU网络中的良性过拟合现象。具体而言,我们考虑标签被相对较小比例破坏或翻转的线性可分数据。我们明确了干净数据间隔的条件,这些条件导致了三种不同的训练结果:良性过拟合,即实现零损失且测试数据以高概率被正确分类;过拟合,即实现零损失但测试数据以常数下界概率被错误分类;以及非过拟合,即干净点(而非损坏点)实现零损失,且测试数据同样以高概率被正确分类。我们的分析提供了训练过程中神经元动态的细致描述,并揭示了两个不同阶段:第一阶段中干净点达到接近零损失,第二阶段中干净点在零损失边界振荡,而损坏点或向零损失收敛或最终被网络归零。我们通过一种组合方法证明了这些结果,该方法涉及对训练这些阶段中干净更新与损坏更新数量的界限分析。