Bilevel optimization has been applied to a wide variety of machine learning models, and numerous stochastic bilevel optimization algorithms have been developed in recent years. However, most existing algorithms restrict their focus on the single-machine setting so that they are incapable of handling the distributed data. To address this issue, under the setting where all participants compose a network and perform peer-to-peer communication in this network, we developed two novel decentralized stochastic bilevel optimization algorithms based on the gradient tracking communication mechanism and two different gradient estimators. Additionally, we established their convergence rates for nonconvex-strongly-convex problems with novel theoretical analysis strategies. To our knowledge, this is the first work achieving these theoretical results. Finally, we applied our algorithms to practical machine learning models, and the experimental results confirmed the efficacy of our algorithms.
翻译:双层优化已广泛应用于多种机器学习模型,近年来也发展出大量随机双层优化算法。然而,现有算法大多局限于单机场景,无法处理分布式数据。为解决这一问题,在所有参与者构成网络并在该网络中进行点对点通信的设定下,我们基于梯度跟踪通信机制和两种不同的梯度估计器,开发了两种新型分布式随机双层优化算法。此外,针对非凸强凸问题,我们采用创新的理论分析策略,建立了算法的收敛速率。据我们所知,这是首个实现这些理论成果的工作。最后,我们将算法应用于实际机器学习模型,实验结果验证了其有效性。