Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimension

from arxiv, Compared to the NeurIPS version (v2), this version strengthens Assumption (K) from d/2<s<=3d/4 to d/2<s<3d/4 and corrects Lemma B.2 by posing additional assumptions. This does not affect any other statements. We provide Python code to reproduce all of our experimental results at https://github.com/moritzhaas/mind-the-spikes

The success of over-parameterized neural networks trained to near-zero training error has caused great interest in the phenomenon of benign overfitting, where estimators are statistically consistent even though they interpolate noisy training data. While benign overfitting in fixed dimension has been established for some learning methods, current literature suggests that for regression with typical kernel methods and wide neural networks, benign overfitting requires a high-dimensional setting where the dimension grows with the sample size. In this paper, we show that the smoothness of the estimators, and not the dimension, is the key: benign overfitting is possible if and only if the estimator's derivatives are large enough. We generalize existing inconsistency results to non-interpolating models and more kernels to show that benign overfitting with moderate derivatives is impossible in fixed dimension. Conversely, we show that rate-optimal benign overfitting is possible for regression with a sequence of spiky-smooth kernels with large derivatives. Using neural tangent kernels, we translate our results to wide neural networks. We prove that while infinite-width networks do not overfit benignly with the ReLU activation, this can be fixed by adding small high-frequency fluctuations to the activation function. Our experiments verify that such neural networks, while overfitting, can indeed generalize well even on low-dimensional data sets.

翻译：在过参数化神经网络训练至接近零训练误差的成功案例中，有益过拟合现象引起了广泛关注——即使估计器对含噪声的训练数据实现了完全拟合，仍能保持统计一致性。虽然固定维度下的有益过拟合已在某些学习方法中得到证实，但现有文献表明，对于典型核方法与宽神经网络的回归任务，有益过拟合需要高维设置，即维度需随样本量增长。本文证明，估计器的平滑性而非维度才是关键：当且仅当估计器的导数足够大时，有益过拟合才可能实现。我们将现有的非一致性结果推广至非插值模型与更多核函数，证明具有适度导数的估计器在固定维度下不可能实现有益过拟合。反之，我们通过构建具有大导数的尖峰-平滑核序列，证明了回归任务中可实现速率最优的有益过拟合。借助神经正切核理论，我们将结论迁移至宽神经网络。研究证明：虽然使用ReLU激活函数的无限宽度网络无法实现有益过拟合，但通过对激活函数添加微小的高频波动即可解决此问题。实验验证表明，此类过拟合神经网络在低维数据集上仍能保持良好的泛化性能。

相关内容

过拟合

关注 8

过拟合，在AI领域多指机器学习得到模型太过复杂，导致在训练集上表现很好，然而在测试集上却不尽人意。过拟合（over-fitting）也称为过学习，它的直观表现是算法在训练集上表现好，但在测试集上表现不好，泛化性能差。过拟合是在模型参数拟合过程中由于训练数据包含抽样误差，在训练时复杂的模型将抽样误差也进行了拟合导致的。

《用于无线通信和传感的智能反射面 (IRS)》（ICC 2022）新加坡国立大学2022最新53页slides

专知会员服务

25+阅读 · 2022年11月16日

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

【ACL2020】多模态信息抽取，365页ppt

专知会员服务

151+阅读 · 2020年7月6日

生成性对抗网络:理论模型、评估指标和最近发展的概述，Generative Adversarial Networks (GANs): An Overview of Theoretical Model, Evaluation Metrics, and Recent Developments

专知会员服务

42+阅读 · 2020年5月30日