How much can you say about the gradient of a neural network without computing a loss or knowing the label? This may sound like a strange question: surely the answer is "very little." However, in this paper, we show that gradients are more structured than previously thought. Gradients lie in a predictable low-dimensional subspace which depends on the network architecture and incoming features. Exploiting this structure can significantly improve gradient-free optimization schemes based on directional derivatives, which have struggled to scale beyond small networks trained on toy datasets. We study how to narrow the gap in optimization performance between methods that calculate exact gradients and those that use directional derivatives. Furthermore, we highlight new challenges in overcoming the large gap between optimizing with exact gradients and guessing the gradients.
翻译:在不计算损失或不了解标签的情况下,你能对神经网络的梯度了解多少?这听起来可能是个奇怪的问题——答案似乎是“非常少”。然而,在本文中,我们证明梯度比先前认为的更具结构性。梯度位于一个可预测的低维子空间中,该子空间取决于网络架构和输入特征。利用这一结构,可以显著改进基于方向导数的无梯度优化方案——这些方案此前难以扩展到在玩具数据集上训练的小型网络之外。我们研究了如何缩小计算精确梯度的方法与使用方向导数的方法之间的优化性能差距。此外,我们强调了在弥合使用精确梯度优化与猜测梯度之间的巨大差距时面临的新挑战。