Bi-level optimization, especially the gradient-based category, has been widely used in the deep learning community including hyperparameter optimization and meta knowledge extraction. Bi-level optimization embeds one problem within another and the gradient-based category solves the outer level task by computing the hypergradient, which is much more efficient than classical methods such as the evolutionary algorithm. In this survey, we first give a formal definition of the gradient-based bi-level optimization. Secondly, we illustrate how to formulate a research problem as a bi-level optimization problem, which is of great practical use for beginners. More specifically, there are two formulations: the single-task formulation to optimize hyperparameters such as regularization parameters and the distilled data, and the multi-task formulation to extract meta knowledge such as the model initialization. With a bi-level formulation, we then discuss four bi-level optimization solvers to update the outer variable including explicit gradient update, proxy update, implicit function update, and closed-form update. Last but not least, we conclude the survey by pointing out the great potential of gradient-based bi-level optimization on science problems (AI4Science).
翻译:双层优化,尤其是基于梯度的类别,已在深度学习社区中广泛应用,包括超参数优化和元知识提取。双层优化将一个子问题嵌入到另一个问题中,基于梯度的类别通过计算超梯度来解决外层任务,其效率远高于进化算法等经典方法。在本综述中,我们首先给出基于梯度的双层优化的正式定义。其次,我们阐述如何将研究问题建模为双层优化问题,这对初学者具有重要的实际意义。具体而言,存在两种建模方式:单任务建模用于优化超参数(如正则化参数和蒸馏数据),以及多任务建模用于提取元知识(如模型初始化)。在建立双层优化模型后,我们讨论了四种用于更新外层变量的双层优化求解器,包括显式梯度更新、代理更新、隐函数更新和闭式更新。最后,我们指出基于梯度的双层优化在科学问题(AI4Science)中的巨大潜力,以此总结本综述。