This paper discusses the estimation of the generalization gap, the difference between generalization performance and training performance, for overparameterized models including neural networks. We first show that a functional variance, a key concept in defining a widely-applicable information criterion, characterizes the generalization gap even in overparameterized settings where a conventional theory cannot be applied. As the computational cost of the functional variance is expensive for the overparameterized models, we propose an efficient approximation of the function variance, the Langevin approximation of the functional variance (Langevin FV). This method leverages only the $1$st-order gradient of the squared loss function, without referencing the $2$nd-order gradient; this ensures that the computation is efficient and the implementation is consistent with gradient-based optimization algorithms. We demonstrate the Langevin FV numerically by estimating the generalization gaps of overparameterized linear regression and non-linear neural network models, containing more than a thousand of parameters therein.
翻译:本文讨论了过参数化模型(包括神经网络)的泛化差距估计问题,即泛化性能与训练性能之差。我们首先证明,泛函方差——这一定义广泛适用信息准则的关键概念——即使在传统理论无法适用的过参数化设置下,也能刻画泛化差距。由于泛函方差在过参数化模型中的计算成本较高,我们提出了一种高效的泛函方差近似方法,即朗之万泛函方差近似(Langevin FV)。该方法仅利用平方损失函数的一阶梯度,无需参考二阶梯度;这确保了计算的高效性,且实现方式与基于梯度的优化算法保持一致。我们通过数值实验展示了Langevin FV的有效性,分别对包含超过一千个参数的过参数化线性回归和非线性神经网络模型的泛化差距进行了估计。