We propose to minimize a generic differentiable loss function with $L_1$ penalty with a redundant reparametrization and straightforward stochastic gradient descent. Our proposal is the direct generalization of a series of previous ideas that the $L_1$ penalty may be equivalent to a differentiable reparametrization with weight decay. We prove that the proposed method, \textit{spred}, is an exact solver of $L_1$ and that the reparametrization trick is completely ``benign" for a generic nonconvex function. Practically, we demonstrate the usefulness of the method in (1) training sparse neural networks to perform gene selection tasks, which involves finding relevant features in a very high dimensional space, and (2) neural network compression task, to which previous attempts at applying the $L_1$-penalty have been unsuccessful. Conceptually, our result bridges the gap between the sparsity in deep learning and conventional statistical learning.
翻译:本文提出通过冗余重参数化结合直接随机梯度下降,对带有$L_1$惩罚项的通用可微损失函数进行最小化。该方案是系列前期工作的直接推广——这些工作表明$L_1$惩罚项等价于带权重衰减的可微重参数化。我们证明所提出的方法(称为\textit{spred})能够精确求解$L_1$范数,且该重参数化技巧对于一般非凸函数具有完全"良性"特性。在实践层面,我们展示了该方法在以下任务中的有效性:(1)训练稀疏神经网络执行基因选择任务(涉及超高维空间相关特征发现);(2)神经网络压缩任务(此前采用$L_1$惩罚项的尝试均未成功)。从概念层面看,该成果弥合了深度学习稀疏性与传统统计学习之间的鸿沟。