This paper investigates efficient deep neural networks (DNNs) to replace dense unstructured weight matrices with structured ones that possess desired properties. The challenge arises because the optimal weight matrix structure in popular neural network models is obscure in most cases and may vary from layer to layer even in the same network. Prior structured matrices proposed for efficient DNNs were mostly hand-crafted without a generalized framework to systematically learn them. To address this issue, we propose a generalized and differentiable framework to learn efficient structures of weight matrices by gradient descent. We first define a new class of structured matrices that covers a wide range of structured matrices in the literature by adjusting the structural parameters. Then, the frequency-domain differentiable parameterization scheme based on the Gaussian-Dirichlet kernel is adopted to learn the structural parameters by proximal gradient descent. Finally, we introduce an effective initialization method for the proposed scheme. Our method learns efficient DNNs with structured matrices, achieving lower complexity and/or higher performance than prior approaches that employ low-rank, block-sparse, or block-low-rank matrices.
翻译:本文研究高效深度神经网络(DNN),以具有理想性质的结构化权重矩阵替代稠密非结构化权重矩阵。挑战在于:主流神经网络模型中最优权重矩阵结构通常不明确,且在同一网络中可能逐层变化。以往面向高效DNN的结构化矩阵多依赖人工设计,缺乏系统学习的通用框架。为解决这一问题,我们提出一种可微的通用框架,通过梯度下降学习权重矩阵的高效结构。首先,我们定义了一类新的结构化矩阵,通过调整结构参数可涵盖文献中多种结构化矩阵。其次,采用基于高斯-狄利克雷核的频域可微参数化方案,通过近端梯度下降学习结构参数。最后,我们为该方案引入有效的初始化方法。本文方法通过结构化矩阵学习高效DNN,在复杂度更低和/或性能更高的条件下,优于采用低秩、块稀疏或块低秩矩阵的现有方法。