We describe the new field of mathematical analysis of deep learning. This field emerged around a list of research questions that were not answered within the classical framework of learning theory. These questions concern: the outstanding generalization power of overparametrized neural networks, the role of depth in deep architectures, the apparent absence of the curse of dimensionality, the surprisingly successful optimization performance despite the non-convexity of the problem, understanding what features are learned, why deep architectures perform exceptionally well in physical problems, and which fine aspects of an architecture affect the behavior of a learning task in which way. We present an overview of modern approaches that yield partial answers to these questions. For selected approaches, we describe the main ideas in more detail.
翻译:我们描述了深度学习数学分析这一新兴领域。该领域围绕一系列在经典学习理论框架内未能解答的研究问题而涌现。这些问题涉及:过度参数化神经网络的卓越泛化能力、深度架构中深度所起的作用、维度诅咒的明显缺失、尽管问题非凸性却出人意料成功的优化性能、理解哪些特征被学习、为何深度架构在物理问题中表现异常出色,以及架构中的哪些细微方面以何种方式影响学习任务的行为。我们概述了能够对这些问题的部分答案提供支撑的现代方法。针对选定的方法,我们更详细地阐述了其主要思想。