This paper presents an approach to addressing the issue of over-parametrization in deep neural networks, more specifically by avoiding the ``sparse double descent'' phenomenon. The authors propose a learning framework that allows avoidance of this phenomenon and improves generalization, an entropy measure to provide more insights on its insurgence, and provide a comprehensive quantitative analysis of various factors such as re-initialization methods, model width and depth, and dataset noise. The proposed approach is supported by experimental results achieved using typical adversarial learning setups. The source code to reproduce the experiments is provided in the supplementary materials and will be publicly released upon acceptance of the paper.
翻译:本文提出了一种解决深度神经网络过参数化问题的方法,具体通过避免“稀疏双下降”现象来实现。作者提出了一种学习框架,能够避免该现象并提升泛化能力;引入了一种熵度量以更深入地揭示该现象的成因;并对重初始化方法、模型宽度与深度、数据集噪声等多种因素进行了全面的定量分析。所提出的方法通过典型对抗学习实验设置下的实验结果得到了验证。重现实验的源代码已提供于补充材料中,将在论文被接收后公开发布。