We study function spaces parametrized by neural networks, referred to as neuromanifolds. Specifically, we focus on deep Multi-Layer Perceptrons (MLPs) and Convolutional Neural Networks (CNNs) with an activation function that is a sufficiently generic polynomial. First, we address the identifiability problem, showing that, for almost all functions in the neuromanifold of an MLP, there exist only finitely many parameter choices yielding that function. For CNNs, the parametrization is generically one-to-one. As a consequence, we compute the dimension of the neuromanifold. Second, we describe singular points of neuromanifolds. We characterize singularities completely for CNNs, and partially for MLPs. In both cases, they arise from sparse subnetworks. For MLPs, we prove that these singularities often correspond to critical points of the mean-squared error loss, which does not hold for CNNs. This provides a geometric explanation of the sparsity bias of MLPs. All of our results leverage tools from algebraic geometry.
翻译:我们研究了由神经网络参数化的函数空间,称为神经流形。具体地,我们关注带有足够一般的多项式激活函数的深度多层感知机(MLP)和卷积神经网络(CNN)。首先,我们探讨辨识性问题,证明对于MLP神经流形中的几乎所有函数,仅有有限个参数选择能够产生该函数。对于CNN,参数化在一般情况下是一一对应的。由此,我们计算了神经流形的维数。其次,我们描述了神经流形的奇异点。我们完整刻画了CNN的奇异性,并部分刻画了MLP的奇异性。在这两种情形中,奇异性都源于稀疏子网络。对于MLP,我们证明这些奇异点通常对应于均方误差损失的临界点,而这一性质对CNN不成立。这为MLP的稀疏偏差提供了几何解释。我们的所有结果均利用了代数几何的工具。