Bilevel optimization has found successful applications in various machine learning problems, including hyper-parameter optimization, data cleaning, and meta-learning. However, its huge computational cost presents a significant challenge for its utilization in large-scale problems. This challenge arises due to the nested structure of the bilevel formulation, where each hyper-gradient computation necessitates a costly inner optimization procedure. To address this issue, we propose a reformulation of bilevel optimization as a minimax problem, effectively decoupling the outer-inner dependency. Under mild conditions, we show these two problems are equivalent. Furthermore, we introduce a multi-stage gradient descent and ascent (GDA) algorithm to solve the resulting minimax problem with convergence guarantees. Extensive experimental results demonstrate that our method outperforms state-of-the-art bilevel methods while significantly reducing the computational cost.
翻译:双层优化已在多种机器学习问题中取得成功应用,包括超参数优化、数据清洗和元学习。然而,其巨大的计算成本对其在大规模问题中的应用构成了重大挑战。这一挑战源于双层公式的嵌套结构,其中每次超梯度计算都需要进行代价高昂的内部优化过程。为解决此问题,我们提出将双层优化重构为极小极大问题,有效解耦了内外层依赖关系。在温和条件下,我们证明这两个问题是等价的。此外,我们引入了一种多阶段梯度下降与上升(GDA)算法来解决由此产生的极小极大问题,并具有收敛性保证。大量实验结果表明,我们的方法在显著降低计算成本的同时,优于最先进的双层优化方法。