This paper investigates the distinctions between gradient methods applied to non-differentiable functions (NGDMs) and classical gradient descents (GDs) designed for differentiable functions. First, we demonstrate significant differences in the convergence properties of NGDMs compared to GDs, challenging the applicability of the extensive neural network convergence literature based on $L-smoothness$ to non-smooth neural networks. Next, we demonstrate the paradoxical nature of NGDM solutions for $L_{1}$-regularized problems, showing that increasing the regularization penalty leads to an increase in the $L_{1}$ norm of optimal solutions in NGDMs. Consequently, we show that widely adopted $L_{1}$ penalization-based techniques for network pruning do not yield expected results. Finally, we explore the Edge of Stability phenomenon, indicating its inapplicability even to Lipschitz continuous convex differentiable functions, leaving its relevance to non-convex non-differentiable neural networks inconclusive. Our analysis exposes misguided interpretations of NGDMs in widely referenced papers and texts due to an overreliance on strong smoothness assumptions, emphasizing the necessity for a nuanced understanding of foundational assumptions in the analysis of these systems.
翻译:本文探讨了应用于非可微函数的梯度方法(NGDMs)与经典梯度下降法(GDs)在可微函数设计上的本质区别。首先,我们展示了NGDMs与GDs在收敛性质上的显著差异,这对基于$L$-光滑性的神经网络收敛性文献在非光滑神经网络中的适用性提出了质疑。其次,我们揭示了NGDMs在$L_{1}$正则化问题中的悖论性质:随着正则化惩罚增强,NGDM最优解的$L_{1}$范数反而增大。由此表明,广泛采用的基于$L_{1}$惩罚的网络剪枝技术无法产生预期效果。最后,我们探讨了稳定性边缘现象,指出该现象甚至不适用于Lipschitz连续凸可微函数,因此其对非凸非可微神经网络的适用性尚无定论。我们的分析揭示了广泛引用文献和教材中因过度依赖强光滑性假设而对NGDMs产生的误读,强调必须深入理解这些系统分析中的基本假设。