While highly expressive parametric models including deep neural networks have an advantage to model complicated concepts, training such highly non-linear models is known to yield a high risk of notorious overfitting. To address this issue, this study considers a $k$th order total variation ($k$-TV) regularization, which is defined as the squared integral of the $k$th order derivative of the parametric models to be trained; penalizing the $k$-TV is expected to yield a smoother function, which is expected to avoid overfitting. While the $k$-TV terms applied to general parametric models are computationally intractable due to the integration, this study provides a stochastic optimization algorithm, that can efficiently train general models with the $k$-TV regularization without conducting explicit numerical integration. The proposed approach can be applied to the training of even deep neural networks whose structure is arbitrary, as it can be implemented by only a simple stochastic gradient descent algorithm and automatic differentiation. Our numerical experiments demonstrate that the neural networks trained with the $K$-TV terms are more ``resilient'' than those with the conventional parameter regularization. The proposed algorithm also can be extended to the physics-informed training of neural networks (PINNs).
翻译:尽管包括深度神经网络在内的高表达能力参数化模型在建模复杂概念方面具有优势,但训练此类高度非线性模型已知会导致严重的过拟合风险。为解决这一问题,本研究考虑了一种$k$阶全变分($k$-TV)正则化,其定义为待训练参数化模型的$k$阶导数的平方积分;对$k$-TV进行惩罚可望获得更平滑的函数,从而避免过拟合。尽管应用于一般参数化模型的$k$-TV项因积分计算而难以处理,但本研究提供了一种随机优化算法,无需进行显式数值积分即可高效训练带有$k$-TV正则化的通用模型。所提方法可应用于训练任意结构的深度神经网络,因为它仅需简单的随机梯度下降算法和自动微分即可实现。数值实验表明,与采用传统参数正则化的神经网络相比,使用$k$-TV项训练的神经网络更具“弹性”。所提算法还可扩展到基于物理信息的神经网络(PINNs)训练中。