This paper introduces a successive affine learning (SAL) model for constructing deep neural networks (DNNs). Traditionally, a DNN is built by solving a non-convex optimization problem. It is often challenging to solve such a problem numerically due to its non-convexity and having a large number of layers. To address this challenge, inspired by the human education system, the multi-grade deep learning (MGDL) model was recently initiated by the author of this paper. The MGDL model learns a DNN in several grades, in each of which one constructs a shallow DNN consisting of a small number of layers. The MGDL model still requires solving several non-convex optimization problems. The proposed SAL model mutates from the MGDL model. Noting that each layer of a DNN consists of an affine map followed by an activation function, we propose to learn the affine map by solving a quadratic/convex optimization problem which involves the activation function only {\it after} the weight matrix and the bias vector for the current layer have been trained. In the context of function approximation, for a given function the SAL model generates an orthogonal expansion of the function with adaptive basis functions in the form of DNNs. We establish the Pythagorean identity and the Parseval identity for the orthogonal system generated by the SAL model. Moreover, we provide a convergence theorem of the SAL process in the sense that either it terminates after a finite number of grades or the norms of its optimal error functions strictly decrease to a limit as the grade number increases to infinity. Furthermore, we present numerical examples of proof of concept which demonstrate that the proposed SAL model significantly outperforms the traditional deep learning model.
翻译:本文提出了一种用于构建深度神经网络(DNN)的逐次仿射学习(SAL)模型。传统上,DNN通过求解非凸优化问题来构建。由于非凸性以及层数众多,此类问题的数值求解往往极具挑战性。受人类教育体系启发,本文作者近期提出了多级深度学习(MGDL)模型以应对这一挑战。该模型分多个阶段学习DNN,每个阶段仅构建包含少量层的浅层网络。然而,MGDL模型仍需求解多个非凸优化问题。本文提出的SAL模型源自MGDL模型。鉴于DNN的每一层均由仿射变换和激活函数组成,我们提出在训练当前层的权重矩阵与偏置向量之后,通过求解一个仅涉及激活函数的二次/凸优化问题来学习仿射变换。在函数逼近的背景下,SAL模型可为给定函数生成以DNN形式表达具有自适应基函数的正交展开。我们建立了SAL模型生成正交系统的勾股恒等式与帕塞瓦尔恒等式,并证明了SAL过程的收敛性定理:该过程要么在有限个阶段后终止,要么随着阶段数趋于无穷,其最优误差函数的范数严格递减至某一极限。此外,我们提供的概念验证数值实验表明,所提出的SAL模型显著优于传统深度学习模型。