Stochastic differential equations (SDEs) have been shown recently to characterize well the dynamics of training machine learning models with SGD. When the generalization error of the SDE approximation closely aligns with that of SGD in expectation, it provides two opportunities for understanding better the generalization behaviour of SGD through its SDE approximation. Firstly, viewing SGD as full-batch gradient descent with Gaussian gradient noise allows us to obtain trajectory-based generalization bound using the information-theoretic bound from Xu and Raginsky [2017]. Secondly, assuming mild conditions, we estimate the steady-state weight distribution of SDE and use information-theoretic bounds from Xu and Raginsky [2017] and Negrea et al. [2019] to establish terminal-state-based generalization bounds. Our proposed bounds have some advantages, notably the trajectory-based bound outperforms results in Wang and Mao [2022], and the terminal-state-based bound exhibits a fast decay rate comparable to stability-based bounds.
翻译:近期研究表明,随机微分方程(SDE)能有效刻画使用随机梯度下降(SGD)训练机器学习模型的动态过程。当SDE近似方法的泛化误差在期望意义上与SGD高度一致时,这为通过SDE近似理解SGD泛化行为提供了两个契机。首先,将SGD视为带有高斯梯度噪声的全批量梯度下降,可借助Xu和Raginsky [2017]的信息论界获得基于轨迹的泛化边界。其次,在温和假设条件下,我们估计SDE的稳态权重分布,并利用Xu和Raginsky [2017]及Negrea等 [2019]的信息论界建立基于终端状态的泛化边界。本文提出的边界具有若干优势,其中基于轨迹的边界优于Wang和Mao [2022]的结果,而基于终端状态的边界展现出可与稳定性边界媲美的快速衰减率。