From Tempered to Benign Overfitting in ReLU Neural Networks

Overparameterized neural networks (NNs) are observed to generalize well even when trained to perfectly fit noisy data. This phenomenon motivated a large body of work on "benign overfitting", where interpolating predictors achieve near-optimal performance. Recently, it was conjectured and empirically observed that the behavior of NNs is often better described as "tempered overfitting", where the performance is non-optimal yet also non-trivial, and degrades as a function of the noise level. However, a theoretical justification of this claim for non-linear NNs has been lacking so far. In this work, we provide several results that aim at bridging these complementing views. We study a simple classification setting with 2-layer ReLU NNs, and prove that under various assumptions, the type of overfitting transitions from tempered in the extreme case of one-dimensional data, to benign in high dimensions. Thus, we show that the input dimension has a crucial role on the type of overfitting in this setting, which we also validate empirically for intermediate dimensions. Overall, our results shed light on the intricate connections between the dimension, sample size, architecture and training algorithm on the one hand, and the type of resulting overfitting on the other hand.

翻译：过参数化的神经网络（NNs）即使在完美拟合含噪数据时，也展现出良好的泛化能力。这一现象催生了大量关于“良性过拟合”的研究——其中插值预测器能够实现接近最优的性能。近期，有学者推测并通过实验观察到，神经网络的行为通常更适合描述为“有节过拟合”，即性能虽非最优但亦非平凡，且会随噪声水平增加而退化。然而，目前尚缺乏对这一论断在非线性神经网络中理论上的合理解释。在本工作中，我们提供了一系列旨在弥合这两种互补视角的研究成果。我们研究了采用双层ReLU神经网络的简单分类设定，并证明在不同假设条件下，过拟合类型会从极端一维数据情形下的有节过拟合，过渡到高维情形下的良性过拟合。因此，我们揭示了输入维度在此设定中对过拟合类型的关键作用，并通过中间维度的实证研究验证了这一结论。总体而言，我们的研究阐明了维度、样本量、架构和训练算法与最终过拟合类型之间错综复杂的关联。

相关内容

过拟合

关注 8

过拟合，在AI领域多指机器学习得到模型太过复杂，导致在训练集上表现很好，然而在测试集上却不尽人意。过拟合（over-fitting）也称为过学习，它的直观表现是算法在训练集上表现好，但在测试集上表现不好，泛化性能差。过拟合是在模型参数拟合过程中由于训练数据包含抽样误差，在训练时复杂的模型将抽样误差也进行了拟合导致的。

UCM《机器学习导论笔记》，80页pdf CSE176 Introduction to Machine Learning

专知会员服务

32+阅读 · 2021年9月29日

语言视觉预训练语言模型揭密，Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models

专知会员服务

36+阅读 · 2020年5月20日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日