Generalization error bounds for deep neural networks trained by stochastic gradient descent (SGD) are derived by combining a dynamical control of an appropriate parameter norm and the Rademacher complexity estimate based on parameter norms. The bounds explicitly depend on the loss along the training trajectory, and work for a wide range of network architectures including multilayer perceptron (MLP) and convolutional neural networks (CNN). Compared with other algorithm-depending generalization estimates such as uniform stability-based bounds, our bounds do not require $L$-smoothness of the nonconvex loss function, and apply directly to SGD instead of Stochastic Langevin gradient descent (SGLD). Numerical results show that our bounds are non-vacuous and robust with the change of optimizer and network hyperparameters.
翻译:通过结合适当参数范数的动态控制与基于参数范数的Rademacher复杂度估计,推导了随机梯度下降训练深度神经网络的泛化误差界。该界限显式依赖于训练轨迹上的损失函数,适用于包括多层感知机和卷积神经网络在内的广泛网络架构。与其他算法依赖型泛化估计(如基于一致稳定性)相比,本文所提界限无需非凸损失函数的L-光滑性假设,且直接适用于SGD而非随机Langevin梯度下降。数值结果表明,该界限非平凡且随优化器与网络超参数的变化保持稳健。