Recent extensive numerical experiments in high scale machine learning have allowed to uncover a quite counterintuitive phase transition, as a function of the ratio between the sample size and the number of parameters in the model. As the number of parameters $p$ approaches the sample size $n$, the generalisation error increases, but surprisingly, it starts decreasing again past the threshold $p=n$. This phenomenon, brought to the theoretical community attention in \cite{belkin2019reconciling}, has been thoroughly investigated lately, more specifically for simpler models than deep neural networks, such as the linear model when the parameter is taken to be the minimum norm solution to the least-squares problem, firstly in the asymptotic regime when $p$ and $n$ tend to infinity, see e.g. \cite{hastie2019surprises}, and recently in the finite dimensional regime and more specifically for linear models \cite{bartlett2020benign}, \cite{tsigler2020benign}, \cite{lecue2022geometrical}. In the present paper, we propose a finite sample analysis of non-linear models of \textit{ridge} type, where we investigate the \textit{overparametrised regime} of the double descent phenomenon for both the \textit{estimation problem} and the \textit{prediction} problem. Our results provide a precise analysis of the distance of the best estimator from the true parameter as well as a generalisation bound which complements recent works of \cite{bartlett2020benign} and \cite{chinot2020benign}. Our analysis is based on tools closely related to the continuous Newton method \cite{neuberger2007continuous} and a refined quantitative analysis of the performance in prediction of the minimum $\ell_2$-norm solution.
翻译:近期大规模机器学习中的大量数值实验揭示了一个相当反直觉的相变现象,其取决于样本量与模型参数数量之比。当参数数量$p$接近样本量$n$时,泛化误差增大,但令人惊讶的是,超过阈值$p=n$后误差又开始下降。这一现象在文献\cite{belkin2019reconciling}中引起了理论界的关注,近期已得到深入研究,特别是针对比深度神经网络更简单的模型,例如当参数取最小范数最小二乘解时的线性模型,首先是在$p$和$n$趋于无穷的渐近情形(参见文献\cite{hastie2019surprises}),最近则是在有限维情形下,尤其针对线性模型(文献\cite{bartlett2020benign}、\cite{tsigler2020benign}、\cite{lecue2022geometrical})。本文中,我们针对岭型非线性模型提出一种有限样本分析,研究双下降现象在过参数化机制下对估计问题和预测问题的影响。我们的结果给出了最优估计量与真实参数之间距离的精确分析,并提供了一个泛化界,补充了文献\cite{bartlett2020benign}和\cite{chinot2020benign}的近期工作。我们的分析基于与连续牛顿法\cite{neuberger2007continuous}密切相关的工具,以及对最小$\ell_2$范数解预测性能的精细化定量分析。