We study the loss landscape of training problems for deep artificial neural networks with a one-dimensional real output whose activation functions contain an affine segment and whose hidden layers have width at least two. It is shown that such problems possess a continuum of spurious (i.e., not globally optimal) local minima for all target functions that are not affine. In contrast to previous works, our analysis covers all sampling and parameterization regimes, general differentiable loss functions, arbitrary continuous nonpolynomial activation functions, and both the finite- and infinite-dimensional setting. It is further shown that the appearance of the spurious local minima in the considered training problems is a direct consequence of the universal approximation theorem and that the underlying mechanisms also cause, e.g., $L^p$-best approximation problems to be ill-posed in the sense of Hadamard for all networks that do not have a dense image. The latter result also holds without the assumption of local affine linearity and without any conditions on the hidden layers.
翻译:我们研究了深度人工神经网络训练问题的损失景观,这类网络具有一维实数输出,其激活函数包含仿射段且隐藏层宽度至少为二。研究表明,对于所有非仿射的目标函数,此类问题均存在连续统的虚假(即非全局最优)局部极小值。与以往工作不同,我们的分析涵盖了所有采样和参数化机制、一般可微损失函数、任意连续非多项式激活函数,以及有限维和无限维设置。进一步证明,所考虑的训练问题中虚假局部极小值的出现是万能逼近定理的直接结果,且其潜在机制也会导致例如:对于所有不具有稠密像的网络,$L^p$-最佳逼近问题在哈达玛意义下是不适定的。后一结论的成立无需局部仿射线性假设,也不依赖于隐藏层的任何条件。