Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimension

The success of over-parameterized neural networks trained to near-zero training error has caused great interest in the phenomenon of benign overfitting, where estimators are statistically consistent even though they interpolate noisy training data. While benign overfitting in fixed dimension has been established for some learning methods, current literature suggests that for regression with typical kernel methods and wide neural networks, benign overfitting requires a high-dimensional setting where the dimension grows with the sample size. In this paper, we show that the smoothness of the estimators, and not the dimension, is the key: benign overfitting is possible if and only if the estimator's derivatives are large enough. We generalize existing inconsistency results to non-interpolating models and more kernels to show that benign overfitting with moderate derivatives is impossible in fixed dimension. Conversely, we show that benign overfitting is possible for regression with a sequence of spiky-smooth kernels with large derivatives. Using neural tangent kernels, we translate our results to wide neural networks. We prove that while infinite-width networks do not overfit benignly with the ReLU activation, this can be fixed by adding small high-frequency fluctuations to the activation function. Our experiments verify that such neural networks, while overfitting, can indeed generalize well even on low-dimensional data sets.

翻译：超参数化神经网络在训练至近乎零误差时取得的成功引起了人们对“良性过拟合”现象的极大兴趣——尽管估计器插值了含噪声的训练数据，它们仍具有统计一致性。尽管对于某些学习方法，已在固定维度下确立了良性过拟合的存在，但当前文献表明，对于使用典型核方法和宽神经网络的回归问题，良性过拟合需要维度随样本量增长的高维设置。本文表明，关键是估计器的光滑性而非维度：当且仅当估计器的导数足够大时，良性过拟合才可能实现。我们将现有的不一致性结果推广至非插值模型及更多核，证明在固定维度下，具有中等导数的良性过拟合是不可能的。反之，我们证明对于具有大导数的一系列尖峰光滑核的回归问题，良性过拟合是可能的。利用神经正切核，我们将结果推广至宽神经网络。我们证明，虽然采用ReLU激活函数的无穷宽网络无法实现良性过拟合，但通过向激活函数添加微小的高频波动可解决此问题。实验验证表明，此类神经网络在过拟合的同时，即使在低维数据集上也能具有良好的泛化能力。

相关内容

过拟合

关注 8

过拟合，在AI领域多指机器学习得到模型太过复杂，导致在训练集上表现很好，然而在测试集上却不尽人意。过拟合（over-fitting）也称为过学习，它的直观表现是算法在训练集上表现好，但在测试集上表现不好，泛化性能差。过拟合是在模型参数拟合过程中由于训练数据包含抽样误差，在训练时复杂的模型将抽样误差也进行了拟合导致的。

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

专知会员服务

54+阅读 · 2021年1月20日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

52+阅读 · 2020年12月14日