Recent studies show that a reproducing kernel Hilbert space (RKHS) is not a suitable space to model functions by neural networks as the curse of dimensionality (CoD) cannot be evaded when trying to approximate even a single ReLU neuron (Bach, 2017). In this paper, we study a suitable function space for over-parameterized two-layer neural networks with bounded norms (e.g., the path norm, the Barron norm) in the perspective of sample complexity and generalization properties. First, we show that the path norm (as well as the Barron norm) is able to obtain width-independence sample complexity bounds, which allows for uniform convergence guarantees. Based on this result, we derive the improved result of metric entropy for $\epsilon$-covering up to $O(\epsilon^{-\frac{2d}{d+2}})$ ($d$ is the input dimension and the depending constant is at most linear order of $d$) via the convex hull technique, which demonstrates the separation with kernel methods with $\Omega(\epsilon^{-d})$ to learn the target function in a Barron space. Second, this metric entropy result allows for building a sharper generalization bound under a general moment hypothesis setting, achieving the rate at $O(n^{-\frac{d+2}{2d+2}})$. Our analysis is novel in that it offers a sharper and refined estimation for metric entropy with a linear dimension dependence and unbounded sampling in the estimation of the sample error and the output error.
翻译:近期研究表明,再生核希尔伯特空间(RKHS)并非模拟神经网络函数的合适空间,因为即使仅尝试逼近单个ReLU神经元时也无法规避维度灾难问题(Bach,2017)。本文从样本复杂度和泛化特性的角度,研究了具有有界范数(如路径范数、Barron范数)的过参数化双层神经网络的适宜函数空间。首先,我们证明路径范数(及Barron范数)能够获得与宽度无关的样本复杂度界限,从而保证一致收敛性。基于此结果,我们通过凸包技术推导出$\epsilon$覆盖度量熵的改进结果$O(\epsilon^{-\frac{2d}{d+2}})$($d$为输入维度,相关常数至多为$d$的线性阶),这揭示了与核方法$\Omega(\epsilon^{-d})$在Barron空间中学习目标函数的本质差异。其次,该度量熵结果使得在一般矩假设条件下能够建立更尖锐的泛化界,达到$O(n^{-\frac{d+2}{2d+2}})$的收敛速率。我们的分析创新之处在于:通过线性维度依赖性和无界采样估计,为样本误差与输出误差的度量熵提供了更精确细致的估计框架。