Understanding when neural networks can be learned efficiently is a fundamental question in learning theory. Existing hardness results suggest that assumptions on both the input distribution and the network's weights are necessary for obtaining efficient algorithms. Moreover, it was previously shown that depth-$2$ networks can be efficiently learned under the assumptions that the input distribution is Gaussian, and the weight matrix is non-degenerate. In this work, we study whether such assumptions may suffice for learning deeper networks and prove negative results. We show that learning depth-$3$ ReLU networks under the Gaussian input distribution is hard even in the smoothed-analysis framework, where a random noise is added to the network's parameters. It implies that learning depth-$3$ ReLU networks under the Gaussian distribution is hard even if the weight matrices are non-degenerate. Moreover, we consider depth-$2$ networks, and show hardness of learning in the smoothed-analysis framework, where both the network parameters and the input distribution are smoothed. Our hardness results are under a well-studied assumption on the existence of local pseudorandom generators.
翻译:理解神经网络何时能被高效学习是学习理论中的一个基本问题。现有的难度结果表明,要获得高效算法,对输入分布和网络权重的假设是必要的。此外,已有研究表明,在输入分布为高斯分布且权重矩阵非退化的假设下,深度为2的网络可以被高效学习。在本研究中,我们探讨这些假设是否足以用于学习更深层的网络,并得到了否定结果。我们证明,即使在平滑分析框架下(即向网络参数添加随机噪声),在高斯输入分布下学习深度为3的ReLU网络仍然是困难的。这意味着,即使权重矩阵非退化,在高斯分布下学习深度为3的ReLU网络也是困难的。此外,我们考虑深度为2的网络,并证明在平滑分析框架下(即网络参数和输入分布都被平滑)学习的难度。我们的难度结果是基于一个被广泛研究的假设,即局部伪随机生成器的存在性。