Sharpness-Aware Minimization (SAM) is a recently proposed gradient-based optimizer (Foret et al., ICLR 2021) that greatly improves the prediction performance of deep neural networks. Consequently, there has been a surge of interest in explaining its empirical success. We focus, in particular, on understanding the role played by normalization, a key component of the SAM updates. We theoretically and empirically study the effect of normalization in SAM for both convex and non-convex functions, revealing two key roles played by normalization: i) it helps in stabilizing the algorithm; and ii) it enables the algorithm to drift along a continuum (manifold) of minima -- a property identified by recent theoretical works that is the key to better performance. We further argue that these two properties of normalization make SAM robust against the choice of hyper-parameters, supporting the practicality of SAM. Our conclusions are backed by various experiments.
翻译:锐度感知极小化(Sharpness-Aware Minimization, SAM)是一种近期提出的基于梯度的优化器(Foret 等,ICLR 2021),它显著提升了深度神经网络的预测性能。因此,解释其经验成功的相关研究激增。我们特别关注理解规范化在SAM更新中扮演的关键角色。我们从理论和实验两方面研究了规范化对凸函数与非凸函数在SAM中的影响,揭示了规范化的两个关键作用:i) 它有助于稳定算法;ii) 它使算法能够沿着极小值连续流形漂移——这一性质被近期理论工作认定为提升性能的关键。我们进一步论证,规范化的这两个特性使SAM对超参数选择具有鲁棒性,从而支撑了SAM的实用性。我们的结论得到了多项实验的支持。