$L_{p}$-norm regularization schemes such as $L_{0}$, $L_{1}$, and $L_{2}$-norm regularization and $L_{p}$-norm-based regularization techniques such as weight decay and group LASSO compute a quantity which depends on model weights considered in isolation from one another. This paper describes a novel regularizer which is not based on an $L_{p}$-norm. In contrast with $L_{p}$-norm-based regularization, this regularizer is concerned with the spatial arrangement of weights within a weight matrix. This regularizer is an additive term for the loss function and is differentiable, simple and fast to compute, scale-invariant, requires a trivial amount of additional memory, and can easily be parallelized. Empirically this method yields approximately a one order-of-magnitude improvement in the number of nonzero model parameters at a given level of accuracy.
翻译:$L_{p}$-范数正则化方案(如$L_{0}$、$L_{1}$和$L_{2}$-范数正则化)以及基于$L_{p}$-范数的正则化技术(如权重衰减和组LASSO)所计算的目标量仅依赖于彼此孤立的模型权重。本文描述了一种新型正则化方法,它并非基于$L_{p}$-范数。与基于$L_{p}$-范数的正则化不同,该正则化方法关注权重矩阵内部权重的空间排列。该正则化方法是损失函数的一个可加项,具有可微性、计算简单快速、尺度不变性、仅需极少量额外内存,且易于并行化的特点。实验结果表明,在给定精度水平下,该方法可使非零模型参数的数量提升约一个数量级。