Sequential-in-time methods solve a sequence of training problems to fit nonlinear parametrizations such as neural networks to approximate solution trajectories of partial differential equations over time. This work shows that sequential-in-time training methods can be understood broadly as either optimize-then-discretize (OtD) or discretize-then-optimize (DtO) schemes, which are well known concepts in numerical analysis. The unifying perspective leads to novel stability and a posteriori error analysis results that provide insights into theoretical and numerical aspects that are inherent to either OtD or DtO schemes such as the tangent space collapse phenomenon, which is a form of over-fitting. Additionally, the unified perspective facilitates establishing connections between variants of sequential-in-time training methods, which is demonstrated by identifying natural gradient descent methods on energy functionals as OtD schemes applied to the corresponding gradient flows.
翻译:时间序列方法通过训练一系列问题来拟合非线性参数化模型(如神经网络),以近似偏微分方程解随时间的演化轨迹。本文表明,时间序列训练方法可广义地理解为优化后离散化(OtD)或离散后优化(DtO)两种策略,这是数值分析中的经典概念。这一统一视角带来了新颖的稳定性和后验误差分析结果,揭示了OtD或DtO方案固有的理论与数值特性(如切空间坍缩现象——一种过拟合形式)。此外,统一视角有助于建立时间序列训练方法变体之间的关联,本文通过识别能量泛函上的自然梯度下降方法作为相应梯度流的OtD方案证明了这一点。