Large foundation models (FMs) adapt surprisingly well to specific domains or tasks with fine-tuning. Federated learning (FL) further enables private FM fine-tuning using the local data on devices. However, the standard FMs' large size poses challenges for resource-constrained and heterogeneous devices. To address this, we consider FMs with reduced parameter sizes, referred to as on-device FMs (ODFMs). While ODFMs allow on-device inference, computational constraints still hinder efficient federated fine-tuning. We propose a parameter-efficient federated fine-tuning method for ODFMs using heterogeneous low-rank approximations (LoRAs) that addresses system and data heterogeneity. We show that homogeneous LoRA ranks face a trade-off between overfitting and slow convergence, and propose HetLoRA, which employs heterogeneous ranks across clients and eliminates the shortcomings of homogeneous HetLoRA. By applying rank self-pruning locally and sparsity-weighted aggregation at the server, we combine the advantages of high and low-rank LoRAs, which achieves improved convergence speed and final performance compared to homogeneous LoRA. Furthermore, it offers enhanced computation efficiency compared to full fine-tuning, making it suitable for heterogeneous devices while preserving data privacy.
翻译:大型基础模型(FMs)在针对特定领域或任务进行微调时展现出惊人的适应性。联邦学习(FL)进一步利用设备上的本地数据实现私有的FM微调。然而,标准FM的巨大规模给资源受限且异构的设备带来了挑战。为解决这一问题,我们考虑采用参数规模缩减的FM,即端侧基础模型(ODFM)。尽管ODFM支持设备端推理,但计算约束仍然阻碍了高效的联邦微调。我们提出了一种面向ODFM的参数高效联邦微调方法,通过利用异构低秩近似(LoRAs)来应对系统与数据的异构性。研究表明,同构LoRA秩面临着过拟合与收敛缓慢之间的权衡,因此我们提出了HetLoRA,该方法在客户端采用异构秩,从而消除了同构HetLoRA的缺陷。通过在本地应用秩自剪枝并在服务器端执行稀疏性加权聚合,我们融合了高秩与低秩LoRA各自的优势,相较于同构LoRA取得了更快的收敛速度和更优的最终性能。此外,与全参数微调相比,该方法显著提升了计算效率,因而适用于异构设备,同时有效保护数据隐私。