When neural networks are trained from data to simulate the dynamics of physical systems, they encounter a persistent challenge: the long-time dynamics they produce are often unphysical or unstable. We analyze the origin of such instabilities when learning linear dynamical systems, focusing on the training dynamics. We make several analytical findings which empirical observations suggest extend to nonlinear dynamical systems. First, the rate of convergence of the training dynamics is uneven and depends on the distribution of energy in the data. As a special case, the dynamics in directions where the data have no energy cannot be learned. Second, in the unlearnable directions, the dynamics produced by the neural network depend on the weight initialization, and common weight initialization schemes can produce unstable dynamics. Third, injecting synthetic noise into the data during training adds damping to the training dynamics and can stabilize the learned simulator, though doing so undesirably biases the learned dynamics. For each contributor to instability, we suggest mitigative strategies. We also highlight important differences between learning discrete-time and continuous-time dynamics, and discuss extensions to nonlinear systems.
翻译:当神经网络通过数据训练来模拟物理系统的动力学时,它们面临一个持续存在的挑战:其产生的长时间动力学行为常常是非物理的或不稳定的。我们分析了在学习线性动力系统时此类不稳定性的起源,重点关注训练动力学过程。我们提出了若干分析性发现,经验观察表明这些发现可推广至非线性动力系统。首先,训练动力学的收敛速率是不均匀的,且取决于数据中的能量分布。作为一种特殊情况,在数据没有能量的方向上,其动力学无法被学习。其次,在不可学习的方向上,神经网络产生的动力学取决于权重初始化,而常见的权重初始化方案可能产生不稳定的动力学。第三,在训练期间向数据注入合成噪声会为训练动力学增加阻尼,从而可能稳定所学习的模拟器,尽管这样做会不可取地使所学动力学产生偏差。针对每种导致不稳定性的因素,我们提出了相应的缓解策略。我们还强调了学习离散时间与连续时间动力学之间的重要差异,并讨论了向非线性系统的扩展。