Works on implicit regularization have studied gradient trajectories during the optimization process to explain why deep networks favor certain kinds of solutions over others. In deep linear networks, it has been shown that gradient descent implicitly regularizes toward low-rank solutions on matrix completion/factorization tasks. Adding depth not only improves performance on these tasks but also acts as an accelerative pre-conditioning that further enhances this bias towards low-rankedness. Inspired by this, we propose an explicit penalty to mirror this implicit bias which only takes effect with certain adaptive gradient optimizers (e.g. Adam). This combination can enable a degenerate single-layer network to achieve low-rank approximations with generalization error comparable to deep linear networks, making depth no longer necessary for learning. The single-layer network also performs competitively or out-performs various approaches for matrix completion over a range of parameter and data regimes despite its simplicity. Together with an optimizer's inductive bias, our findings suggest that explicit regularization can play a role in designing different, desirable forms of regularization and that a more nuanced understanding of this interplay may be necessary.
翻译:关于隐式正则化的研究通过分析优化过程中的梯度轨迹,解释了为何深度网络倾向于特定类型的解。在深度线性网络中,梯度下降被证明能在矩阵补全/分解任务中隐式正则化至低秩解。增加网络深度不仅提升了这些任务的性能,还起到加速预处理作用,进一步增强了低秩偏好。受此启发,我们提出一种显式惩罚项以镜像此隐式偏差,该惩罚项仅在使用特定自适应梯度优化器(如Adam)时生效。这种结合能使退化的单层网络实现与深度线性网络相当泛化误差的低秩近似,从而使得深度不再成为学习的必要条件。尽管结构简单,该单层网络在多种参数与数据规模下的矩阵补全任务中,性能可与各类先进方法媲美甚至超越。结合优化器的归纳偏差,我们的发现表明:显式正则化可在设计不同形式的期望正则化中发挥作用,而对此类交互机制的更精细理解或许不可或缺。