We examine gradient descent on unregularized logistic regression problems, with homogeneous linear predictors on linearly separable datasets. We show the predictor converges to the direction of the max-margin (hard margin SVM) solution. The result also generalizes to other monotone decreasing loss functions with an infimum at infinity, to multi-class problems, and to training a weight layer in a deep network in a certain restricted setting. Furthermore, we show this convergence is very slow, and only logarithmic in the convergence of the loss itself. This can help explain the benefit of continuing to optimize the logistic or cross-entropy loss even after the training error is zero and the training loss is extremely small, and, as we show, even if the validation loss increases. Our methodology can also aid in understanding implicit regularization n more complex models and with other optimization methods.
翻译:我们研究了线性可分数据集上同质线性预测器的无正则化逻辑回归问题中的梯度下降行为。结果表明,预测器收敛至最大间隔(硬间隔SVM)解的方向。该结论还可推广至其他具有无穷远处下确界的单调递减损失函数、多类问题,以及在特定受限设置下深度网络中权重层的训练。此外,我们证明这种收敛速度非常缓慢,仅表现为损失函数本身的对数收敛性。这有助于解释为何在训练误差为零且训练损失极小时,仍继续优化逻辑损失或交叉熵损失的收益,同时我们证明即使在验证损失上升的情况下也是如此。我们的方法论也有助于理解更复杂模型及其他优化方法中的隐式正则化现象。