In recent years, we have witnessed the emergence of scientific machine learning as a data-driven tool for the analysis, by means of deep-learning techniques, of data produced by computational science and engineering applications. At the core of these methods is the supervised training algorithm to learn the neural network realization, a highly non-convex optimization problem that is usually solved using stochastic gradient methods. However, distinct from deep-learning practice, scientific machine-learning training problems feature a much larger volume of smooth data and better characterizations of the empirical risk functions, which make them suited for conventional solvers for unconstrained optimization. We introduce a lightweight software framework built on top of the Portable and Extensible Toolkit for Scientific computation to bridge the gap between deep-learning software and conventional solvers for unconstrained minimization. We empirically demonstrate the superior efficacy of a trust region method based on the Gauss-Newton approximation of the Hessian in improving the generalization errors arising from regression tasks when learning surrogate models for a wide range of scientific machine-learning techniques and test cases. All the conventional second-order solvers tested, including L-BFGS and inexact Newton with line-search, compare favorably, either in terms of cost or accuracy, with the adaptive first-order methods used to validate the surrogate models.
翻译:近年来,我们见证了科学机器学习作为一种数据驱动工具的出现,通过深度学习技术,用于分析计算科学与工程应用所生成的数据。这类方法的核心是监督训练算法,用于学习神经网络的实现——这是一个高度非凸的优化问题,通常采用随机梯度方法求解。然而,与深度学习实践不同,科学机器学习训练问题具有更大规模的平滑数据和更清晰的经验风险函数表征,使其适用于传统的无约束优化求解器。我们引入了一个轻量级软件框架,基于可移植可扩展科学计算工具包构建,以弥合深度学习软件与无约束最小化传统求解器之间的差距。我们通过实验证明,基于高斯-牛顿海森矩阵近似的信赖域方法在降低回归任务泛化误差方面具有显著优势,该方法适用于学习多种科学机器学习技术与测试案例的代理模型。所有测试的传统二阶求解器(包括L-BFGS和非精确牛顿法配合线搜索)与用于验证代理模型的自适应一阶方法相比,在计算成本或精度方面均表现更优。