Study of the nonlinear evolution deep neural network (DNN) parameters undergo during training has uncovered regimes of distinct dynamical behavior. While a detailed understanding of these phenomena has the potential to advance improvements in training efficiency and robustness, the lack of methods for identifying when DNN models have equivalent dynamics limits the insight that can be gained from prior work. Topological conjugacy, a notion from dynamical systems theory, provides a precise definition of dynamical equivalence, offering a possible route to address this need. However, topological conjugacies have historically been challenging to compute. By leveraging advances in Koopman operator theory, we develop a framework for identifying conjugate and non-conjugate training dynamics. To validate our approach, we demonstrate that it can correctly identify a known equivalence between online mirror descent and online gradient descent. We then utilize it to: identify non-conjugate training dynamics between shallow and wide fully connected neural networks; characterize the early phase of training dynamics in convolutional neural networks; uncover non-conjugate training dynamics in Transformers that do and do not undergo grokking. Our results, across a range of DNN architectures, illustrate the flexibility of our framework and highlight its potential for shedding new light on training dynamics.
翻译:对深度神经网络参数在训练过程中经历的非线性演化研究揭示了不同动力学行为的机制。尽管对这些现象的深入理解有望推动训练效率与鲁棒性的改进,但缺乏识别DNN模型何时具有等效动力学的方法限制了从先前工作中获得的洞见。拓扑共轭性——源自动力系统理论的概念——为动力学等效性提供了精确定义,为满足这一需求提供了可能途径。然而,拓扑共轭性的计算历来具有挑战性。通过利用Koopman算子理论的进展,我们开发了一个识别共轭与非共轭训练动力学的框架。为验证该方法,我们证明其能正确识别在线镜像下降法与在线梯度下降法之间的已知等效性。随后我们运用该框架:识别浅层与宽全连接神经网络之间的非共轭训练动力学;表征卷积神经网络训练早期的动力学特征;揭示Transformer模型中经历与未经历顿悟现象的非共轭训练动力学。我们在多种DNN架构上获得的结果,展示了该框架的灵活性,并凸显了其为训练动力学研究提供新见解的潜力。