We establish matching upper and lower generalization error bounds for mini-batch Gradient Descent (GD) training with either deterministic or stochastic, data-independent, but otherwise arbitrary batch selection rules. We consider smooth Lipschitz-convex/nonconvex/strongly-convex loss functions, and show that classical upper bounds for Stochastic GD (SGD) also hold verbatim for such arbitrary nonadaptive batch schedules, including all deterministic ones. Further, for convex and strongly-convex losses we prove matching lower bounds directly on the generalization error uniform over the aforementioned class of batch schedules, showing that all such batch schedules generalize optimally. Lastly, for smooth (non-Lipschitz) nonconvex losses, we show that full-batch (deterministic) GD is essentially optimal, among all possible batch schedules within the considered class, including all stochastic ones.
翻译:我们针对使用确定性或随机、数据独立但任意批次选择规则的小批量梯度下降训练,建立了匹配的泛化误差上下界。我们考虑了光滑利普希茨凸/非凸/强凸损失函数,并证明随机梯度下降的经典上界对这些任意非自适应批次调度(包括所有确定性调度)同样直接成立。进一步地,对于凸和强凸损失,我们直接针对上述批次调度类别上的泛化误差证明了匹配的下界,表明所有此类批次调度均能实现最优泛化。最后,对于光滑(非利普希茨)非凸损失,我们证明在所考虑类别内的所有批次调度(包括所有随机调度)中,全批量(确定性)梯度下降本质上是最优的。