We analyze VeLO (versatile learned optimizer), the largest scale attempt to train a general purpose "foundational" optimizer to date. VeLO was trained on thousands of machine learning tasks using over 4000 TPU months with the goal of producing an optimizer capable of generalizing to new problems while being hyperparameter free, and outperforming industry standards such as Adam. We independently evaluate VeLO on the MLCommons optimizer benchmark suite. We find that, contrary to initial claims: (1) VeLO has a critical hyperparameter that needs problem-specific tuning, (2) VeLO does not necessarily outperform competitors in quality of solution found, and (3) VeLO is not faster than competing optimizers at reducing the training loss. These observations call into question VeLO's generality and the value of the investment in training it.
翻译:我们分析了VeLO(通用型学习优化器),这是迄今最大规模训练通用“基础”优化器的尝试。VeLO使用超过4000 TPU月在数千个机器学习任务上进行训练,旨在生成一种无需超参数即可泛化到新问题、并超越Adam等行业标准的优化器。我们在MLCommons优化器基准测试套件上独立评估了VeLO。研究发现,与初步宣称相反:(1)VeLO存在需要针对具体问题调整的关键超参数;(2)VeLO在解质量上未必胜过竞争对手;(3)VeLO在降低训练损失方面并不比竞争优化器更快。这些观察结果对VeLO的通用性及其训练投入的价值提出了质疑。