This paper studies the expressive power of deep neural networks from the perspective of function compositions. We show that repeated compositions of a single fixed-size ReLU network can produce super expressive power. In particular, we prove by construction that $\mathcal{L}_2\circ \boldsymbol{g}^{\circ r}\circ \boldsymbol{\mathcal{L}}_1$ can approximate $1$-Lipschitz continuous functions on $[0,1]^d$ with an error $\mathcal{O}(r^{-1/d})$, where $\boldsymbol{g}$ is realized by a fixed-size ReLU network, $\boldsymbol{\mathcal{L}}_1$ and $\mathcal{L}_2$ are two affine linear maps matching the dimensions, and $\boldsymbol{g}^{\circ r}$ means the $r$-times composition of $\boldsymbol{g}$. Furthermore, we extend such a result to generic continuous functions on $[0,1]^d$ with the approximation error characterized by the modulus of continuity. Our results reveal that a continuous-depth network generated via a dynamical system has good approximation power even if its dynamics function is time-independent and realized by a fixed-size ReLU network.
翻译:本文从函数复合的角度研究深度神经网络的表达能力。我们证明了单一固定大小ReLU网络的重复复合可以产生极强的表达能力。具体而言,通过构造性证明,$\mathcal{L}_2\circ \boldsymbol{g}^{\circ r}\circ \boldsymbol{\mathcal{L}}_1$ 能以误差 $\mathcal{O}(r^{-1/d})$ 逼近 $[0,1]^d$ 上的 $1$-Lipschitz 连续函数,其中 $\boldsymbol{g}$ 由固定大小的ReLU网络实现,$\boldsymbol{\mathcal{L}}_1$ 和 $\mathcal{L}_2$ 是两个维度匹配的仿射线性映射,而 $\boldsymbol{g}^{\circ r}$ 表示 $\boldsymbol{g}$ 的 $r$ 次复合。进一步地,我们将该结果推广到 $[0,1]^d$ 上的通用连续函数,其逼近误差由连续模刻画。我们的结果表明,即使动力系统生成的深度网络其动力学函数是时不变的,且由固定大小的ReLU网络实现,该网络仍具有良好的逼近能力。