Given a training set, a loss function, and a neural network architecture, it is often taken for granted that optimal network parameters exist, and a common practice is to apply available optimization algorithms to search for them. In this work, we show that the existence of an optimal solution is not always guaranteed, especially in the context of {\em sparse} ReLU neural networks. In particular, we first show that optimization problems involving deep networks with certain sparsity patterns do not always have optimal parameters, and that optimization algorithms may then diverge. Via a new topological relation between sparse ReLU neural networks and their linear counterparts, we derive -- using existing tools from real algebraic geometry -- an algorithm to verify that a given sparsity pattern suffers from this issue. Then, the existence of a global optimum is proved for every concrete optimization problem involving a shallow sparse ReLU neural network of output dimension one. Overall, the analysis is based on the investigation of two topological properties of the space of functions implementable as sparse ReLU neural networks: a best approximation property, and a closedness property, both in the uniform norm. This is studied both for (finite) domains corresponding to practical training on finite training sets, and for more general domains such as the unit cube. This allows us to provide conditions for the guaranteed existence of an optimum given a sparsity pattern. The results apply not only to several sparsity patterns proposed in recent works on network pruning/sparsification, but also to classical dense neural networks, including architectures not covered by existing results.
翻译:给定训练集、损失函数和神经网络架构,通常假定最优网络参数存在,且常见做法是应用现有优化算法进行搜索。本文表明,最优解的存在性并非总能得到保证,尤其是在稀疏ReLU神经网络中。具体而言,我们首先证明,具有特定稀疏模式的深度网络优化问题并非总是存在最优参数,且此时优化算法可能发散。通过建立稀疏ReLU神经网络与其线性对应物之间的新拓扑关系,我们利用实代数几何的现有工具推导出一种算法,用于验证给定稀疏模式是否存在此问题。接着,我们证明了每个涉及输出维度为一的浅层稀疏ReLU神经网络的具体优化问题均存在全局最优解。总体而言,该分析基于对稀疏ReLU神经网络可实现的函数空间两种拓扑性质的探究:最佳逼近性质与闭包性质(两者均在一致范数下)。我们既针对有限训练集上的实际训练所对应的有限定义域,也针对更一般的定义域(如单位立方体)研究了该问题。这使我们能够为给定稀疏模式下最优解的可保证存在性提供条件。这些结果不仅适用于近期网络剪枝/稀疏化研究中提出的多种稀疏模式,还适用于经典稠密神经网络,包括现有结果未覆盖的架构。