The confluence of Federated Learning (FL) and Large Language Models (LLMs) is ushering in a new era in privacy-preserving natural language processing. However, the intensive memory requirements for fine-tuning LLMs pose significant challenges, especially when deploying on clients with limited computational resources. To circumvent this, we explore the novel integration of Memory-efficient Zeroth-Order Optimization within a federated setting, a synergy we term as FedMeZO. Our study is the first to examine the theoretical underpinnings of FedMeZO in the context of LLMs, tackling key questions regarding the influence of large parameter spaces on optimization behavior, the establishment of convergence properties, and the identification of critical parameters for convergence to inform personalized federated strategies. Our extensive empirical evidence supports the theory, showing that FedMeZO not only converges faster than traditional first-order methods such as FedAvg but also significantly reduces GPU memory usage during training to levels comparable to those during inference. Moreover, the proposed personalized FL strategy that is built upon the theoretical insights to customize the client-wise learning rate can effectively accelerate loss reduction. We hope our work can help to bridge theoretical and practical aspects of federated fine-tuning for LLMs, thereby stimulating further advancements and research in this area.
翻译:联邦学习(FL)与大语言模型(LLMs)的融合正在开启隐私保护自然语言处理的新纪元。然而,微调LLMs所需的大量内存带来了重大挑战,尤其是在计算资源有限的客户端上部署时。为规避此问题,我们探索了内存高效的零阶优化在联邦环境中的新颖集成,这一协同作用我们称之为FedMeZO。本研究首次在LLMs背景下检验FedMeZO的理论基础,解决了关于大规模参数空间对优化行为的影响、收敛性质的建立以及识别收敛关键参数以指导个性化联邦策略等关键问题。我们广泛的实证证据支持该理论,表明FedMeZO不仅比传统一阶方法(如FedAvg)收敛更快,还能在训练期间将GPU内存使用显著降低至与推理期间相当的水平。此外,基于理论见解构建的个性化联邦学习策略,通过定制客户端特定的学习率,能有效加速损失下降。我们希望本工作有助于弥合LLMs联邦微调的理论与实践,从而推动该领域的进一步发展和研究。