Unraveling the reasons behind the remarkable success and exceptional generalization capabilities of deep neural networks presents a formidable challenge. Recent insights from random matrix theory, specifically those concerning the spectral analysis of weight matrices in deep neural networks, offer valuable clues to address this issue. A key finding indicates that the generalization performance of a neural network is associated with the degree of heavy tails in the spectrum of its weight matrices. To capitalize on this discovery, we introduce a novel regularization technique, termed Heavy-Tailed Regularization, which explicitly promotes a more heavy-tailed spectrum in the weight matrix through regularization. Firstly, we employ the Weighted Alpha and Stable Rank as penalty terms, both of which are differentiable, enabling the direct calculation of their gradients. To circumvent over-regularization, we introduce two variations of the penalty function. Then, adopting a Bayesian statistics perspective and leveraging knowledge from random matrices, we develop two novel heavy-tailed regularization methods, utilizing Powerlaw distribution and Frechet distribution as priors for the global spectrum and maximum eigenvalues, respectively. We empirically show that heavytailed regularization outperforms conventional regularization techniques in terms of generalization performance.
翻译:揭示深度神经网络卓越成功与非凡泛化能力背后的原因是一项艰巨挑战。近期随机矩阵理论的研究进展,特别是关于深度神经网络权重矩阵谱分析的发现,为破解这一难题提供了宝贵线索。关键结果表明,神经网络泛化性能与其权重矩阵谱的重尾程度存在关联。为利用这一发现,我们提出一种名为重尾正则化的新型正则化技术,通过显式促进权重矩阵谱呈现更强的重尾特性来实现。首先,我们采用加权α指数和稳定秩作为惩罚项,二者均为可微函数,可直接计算其梯度。为避免过度正则化,我们引入两种惩罚函数变体。随后,从贝叶斯统计视角出发,结合随机矩阵理论成果,我们开发了两种新型重尾正则化方法,分别采用幂律分布和弗雷歇分布作为全局谱与最大特征值的先验分布。实验证明,重尾正则化在泛化性能上优于传统正则化技术。