Hyperparameter tuning, particularly the selection of an appropriate learning rate in adaptive gradient training methods, remains a challenge. To tackle this challenge, in this paper, we propose a novel parameter-free optimizer, AdamG (Adam with the golden step size), designed to automatically adapt to diverse optimization problems without manual tuning. The core technique underlying AdamG is our golden step size derived for the AdaGrad-Norm algorithm, which is expected to help AdaGrad-Norm preserve the tuning-free convergence and approximate the optimal step size in expectation w.r.t. various optimization scenarios. To better evaluate tuning-free performance, we propose a novel evaluation criterion, stability, to comprehensively assess the efficacy of parameter-free optimizers in addition to classical performance criteria. Empirical results demonstrate that compared with other parameter-free baselines, AdamG achieves superior performance, which is consistently on par with Adam using a manually tuned learning rate across various optimization tasks.
翻译:超参数调优,特别是自适应梯度训练方法中适当学习率的选择,仍然是一大挑战。为解决这一问题,本文提出了一种新颖的无参数优化器AdamG(采用黄金步长的Adam),旨在无需手动调优的情况下自动适应各类优化问题。AdamG的核心技术在于我们为AdaGrad-Norm算法推导的黄金步长,该步长有望帮助AdaGrad-Norm保持免调优收敛性,并在各种优化场景下近似期望中的最优步长。为更全面地评估免调优性能,除传统性能指标外,我们提出了一种新的评估准则——稳定性,以综合衡量无参数优化器的有效性。实验结果表明,与其他无参数基线方法相比,AdamG实现了卓越的性能,在各种优化任务中始终与使用手动调优学习率的Adam表现相当。