Modern deep learning models with great expressive power can be trained to overfit the training data but still generalize well. This phenomenon is referred to as \textit{benign overfitting}. Recently, a few studies have attempted to theoretically understand benign overfitting in neural networks. However, these works are either limited to neural networks with smooth activation functions or to the neural tangent kernel regime. How and when benign overfitting can occur in ReLU neural networks remains an open problem. In this work, we seek to answer this question by establishing algorithm-dependent risk bounds for learning two-layer ReLU convolutional neural networks with label-flipping noise. We show that, under mild conditions, the neural network trained by gradient descent can achieve near-zero training loss and Bayes optimal test risk. Our result also reveals a sharp transition between benign and harmful overfitting under different conditions on data distribution in terms of test risk. Experiments on synthetic data back up our theory.
翻译:现代深度学习模型具有强大的表达能力,可以在训练数据上过拟合但仍然具有良好的泛化性能,这一现象被称为“良性过拟合”。近期,一些研究尝试从理论上理解神经网络中的良性过拟合,但这些工作要么局限于使用平滑激活函数的神经网络,要么局限于神经正切核框架。在ReLU神经网络中良性过拟合如何发生及何时发生仍是一个开放性问题。本研究通过建立带有标签翻转噪声的双层ReLU卷积神经网络的学习算法依赖风险界,寻求回答这一问题。我们证明,在温和条件下,由梯度下降训练的神经网络能够实现近乎为零的训练损失和贝叶斯最优测试风险。我们的结果还揭示了在数据分布不同条件下,测试风险从良性过拟合向有害过拟合的剧烈转变。在合成数据上的实验支持了我们的理论。