This paper introduces a novel theoretical framework for the analysis of vector-valued neural networks through the development of vector-valued variation spaces, a new class of reproducing kernel Banach spaces. These spaces emerge from studying the regularization effect of weight decay in training networks with activations like the rectified linear unit (ReLU). This framework offers a deeper understanding of multi-output networks and their function-space characteristics. A key contribution of this work is the development of a representer theorem for the vector-valued variation spaces. This representer theorem establishes that shallow vector-valued neural networks are the solutions to data-fitting problems over these infinite-dimensional spaces, where the network widths are bounded by the square of the number of training data. This observation reveals that the norm associated with these vector-valued variation spaces encourages the learning of features that are useful for multiple tasks, shedding new light on multi-task learning with neural networks. Finally, this paper develops a connection between weight-decay regularization and the multi-task lasso problem. This connection leads to novel bounds for layer widths in deep networks that depend on the intrinsic dimensions of the training data representations. This insight not only deepens the understanding of the deep network architectural requirements, but also yields a simple convex optimization method for deep neural network compression. The performance of this compression procedure is evaluated on various architectures.
翻译:本文通过构建向量值变分空间(一类新的再生核巴拿赫空间),提出了分析向量值神经网络的新理论框架。该框架源于研究使用修正线性单元(ReLU)等激活函数训练网络时权重衰减的正则化效应,为深入理解多输出网络及其函数空间特性提供了新视角。本文的核心贡献在于建立了向量值变分空间的表示定理,该定理表明浅层向量值神经网络是这些无限维空间上数据拟合问题的解,且网络宽度受限于训练数据数量的平方。这一发现揭示了与这些向量值变分空间相关联的范数鼓励学习多任务共享特征,为多任务学习提供了新见解。最后,本文建立了权重衰减正则化与多任务套索问题的联系,由此推导出依赖训练数据表示本征维度的深度网络层宽新边界。该洞见不仅深化了对深度网络架构需求的理解,还衍生出一种用于深度神经网络压缩的简单凸优化方法,并在多种架构上评估了该压缩策略的性能。