We consider the optimization problem associated with fitting two-layer ReLU networks with respect to the squared loss, where labels are assumed to be generated by a target network. Focusing first on standard Gaussian inputs, we show that the structure of spurious local minima detected by stochastic gradient descent (SGD) is, in a well-defined sense, the \emph{least loss of symmetry} with respect to the target weights. A closer look at the analysis indicates that this principle of least symmetry breaking may apply to a broader range of settings. Motivated by this, we conduct a series of experiments which corroborate this hypothesis for different classes of non-isotropic non-product distributions, smooth activation functions and networks with a few layers.
翻译:我们研究了与平方损失相关的两层ReLU网络拟合优化问题,其中标签假定由目标网络生成。首先聚焦于标准高斯输入,我们证明了随机梯度下降(SGD)检测到的虚假局部极小值的结构,在严格定义的意义上,相对于目标权重表现为*最小对称性损失*。进一步分析表明,这一最小对称破缺原理可能适用于更广泛的场景。受此启发,我们开展了一系列实验,验证了该假设在不同类别的非各向同性非乘积分布、平滑激活函数以及少量层网络中的有效性。