We show that gating mechanisms in recurrent neural networks (RNNs) induce lag-dependent and direction-dependent effective learning rates, even when training uses a fixed, global step size. This behavior arises from a coupling between state-space time-scales (parametrized by the gates) and parameter-space dynamics during gradient descent. By deriving exact Jacobians for leaky-integrator and gated RNNs and applying a first-order expansion, we make explicit how constant, scalar, and multi-dimensional gates reshape gradient propagation, modulate effective step sizes, and introduce anisotropy in parameter updates. These findings reveal that gates act not only as filters of information flow, but also as data-driven preconditioners of optimization, with formal connections to learning-rate schedules, momentum, and adaptive methods such as Adam. Empirical simulations corroborate these predictions: across several sequence tasks, gates produce lag-dependent effective learning rates and concentrate gradient flow into low-dimensional subspaces, matching or exceeding the anisotropic structure induced by Adam. Notably, gating and optimizer-driven adaptivity shape complementary aspects of credit assignment: gates align state-space transport with loss-relevant directions, while optimizers rescale parameter-space updates. Overall, this work provides a unified dynamical systems perspective on how gating couples state evolution with parameter updates, clarifying why gated architectures achieve robust trainability in practice.
翻译:我们证明,即使使用固定的全局步长进行训练,循环神经网络(RNN)中的门控机制也会诱导出依赖于滞后和方向的有效学习率。这种行为的根源在于梯度下降过程中状态空间时间尺度(由门控参数化)与参数空间动力学之间的耦合。通过推导漏积分器和门控RNN的精确雅可比矩阵,并应用一阶展开,我们明确了常数门、标量门和多维门如何重塑梯度传播、调节有效步长,并引入参数更新的各向异性。这些发现揭示门控不仅充当信息流的滤波器,还作为优化过程中的数据驱动预处理器,与学习率调度、动量方法和Adam等自适应方法存在形式上的联系。实验模拟验证了这些预测:在多个序列任务中,门控产生滞后依赖的有效学习率,并将梯度流集中于低维子空间,其性能达到或超过Adam所诱导的各向异性结构。值得注意的是,门控与优化器驱动的自适应性在信用分配中塑造互补方面:门控将状态空间传输沿损失相关方向对齐,而优化器则重新缩放参数空间更新。总体而言,本研究为门控如何耦合状态演化与参数更新提供了统一的动力系统视角,阐明了门控架构在实践中实现稳健可训练性的内在机制。