Deep neural networks have been demonstrated to achieve phenomenal success in many domains, and yet their inner mechanisms are not well understood. In this paper, we investigate the curvature of image manifolds, i.e., the manifold deviation from being flat in its principal directions. We find that state-of-the-art trained convolutional neural networks for image classification have a characteristic curvature profile along layers: an initial steep increase, followed by a long phase of a plateau, and followed by another increase. In contrast, this behavior does not appear in untrained networks in which the curvature flattens. We also show that the curvature gap between the last two layers has a strong correlation with the generalization capability of the network. Moreover, we find that the intrinsic dimension of latent codes is not necessarily indicative of curvature. Finally, we observe that common regularization methods such as mixup yield flatter representations when compared to other methods. Our experiments show consistent results over a variety of deep learning architectures and multiple data sets. Our code is publicly available at https://github.com/azencot-group/CRLM
翻译:深度神经网络已在众多领域取得显著成功,但其内部机制尚未得到充分理解。本文研究图像流形的曲率,即流形在主方向上偏离平坦的程度。我们发现,用于图像分类的最先进卷积神经网络沿各层呈现出特征性曲率分布:初始急剧上升,随后是长期平台期,最后再次上升。相比之下,未训练网络中不存在这种曲率平坦化的行为。我们还表明,最后两层之间的曲率差距与网络的泛化能力密切相关。此外,我们发现潜在编码的内在维度不一定能指示曲率。最后,我们观察到与其它方法相比,mixup等常用正则化方法能产生更平坦的表示。我们的实验在多种深度学习架构和数据集上得到了一致的结果。代码已公开于https://github.com/azencot-group/CRLM。