In practice, deep neural networks are often able to easily interpolate their training data. To understand this phenomenon, many works have aimed to quantify the memorization capacity of a neural network architecture: the largest number of points such that the architecture can interpolate any placement of these points with any assignment of labels. For real-world data, however, one intuitively expects the presence of a benign structure so that interpolation already occurs at a smaller network size than suggested by memorization capacity. In this paper, we investigate interpolation by adopting an instance-specific viewpoint. We introduce a simple randomized algorithm that, given a fixed finite dataset with two classes, with high probability constructs an interpolating three-layer neural network in polynomial time. The required number of parameters is linked to geometric properties of the two classes and their mutual arrangement. As a result, we obtain guarantees that are independent of the number of samples and hence move beyond worst-case memorization capacity bounds. We illustrate the effectiveness of the algorithm in non-pathological situations with extensive numerical experiments and link the insights back to the theoretical results.
翻译:在实践中,深度神经网络通常能够轻松地内插其训练数据。为理解这一现象,许多研究致力于量化神经网络架构的记忆容量:即该架构在任意数据点放置与标签分配下仍能内插的最大样本数量。然而,对于现实世界数据,人们直觉上预期良性结构的存在,使得内插在比记忆容量所指示的更小网络规模下即可发生。本文从实例化视角出发研究内插问题。我们提出一种简单的随机化算法,对于给定的固定二类有限数据集,该算法能以高概率在多项式时间内构建一个三层内插神经网络。所需参数数量与两个类别的几何特性及其相互排列有关。由此,我们获得了与样本数量无关的保证,从而超越最坏情况下的记忆容量界限。我们通过大量数值实验展示了该算法在非病态场景中的有效性,并将相关洞见与理论结果相联系。