Federated learning (FL) has been widely adopted for collaborative training on decentralized data. However, it faces the challenges of data, system, and model heterogeneity. This has inspired the emergence of model-heterogeneous personalized federated learning (MHPFL). Nevertheless, the problem of ensuring data and model privacy, while achieving good model performance and keeping communication and computation costs low remains open in MHPFL. To address this problem, we propose a model-heterogeneous personalized Federated learning with Mixture of Experts (pFedMoE) method. It assigns a shared homogeneous small feature extractor and a local gating network for each client's local heterogeneous large model. Firstly, during local training, the local heterogeneous model's feature extractor acts as a local expert for personalized feature (representation) extraction, while the shared homogeneous small feature extractor serves as a global expert for generalized feature extraction. The local gating network produces personalized weights for extracted representations from both experts on each data sample. The three models form a local heterogeneous MoE. The weighted mixed representation fuses generalized and personalized features and is processed by the local heterogeneous large model's header with personalized prediction information. The MoE and prediction header are updated simultaneously. Secondly, the trained local homogeneous small feature extractors are sent to the server for cross-client information fusion via aggregation. Overall, pFedMoE enhances local model personalization at a fine-grained data level, while supporting model heterogeneity.
翻译:联邦学习(FL)已被广泛用于分散数据的协同训练。然而,它面临数据、系统和模型异构性的挑战。这激发了模型异构个性化联邦学习(MHPFL)的出现。尽管如此,如何在MHPFL中确保数据和模型隐私,同时实现良好的模型性能并保持较低的通信和计算成本,仍然是一个开放性问题。为解决此问题,我们提出了一种基于专家混合的模型异构个性化联邦学习方法(pFedMoE)。该方法为每个客户端为局部异构大模型配备一个共享的同构小型特征提取器和一个局部门控网络。首先,在局部训练中,局部异构模型的特征提取器充当局部专家,用于个性化特征(表示)提取,而共享的同构小型特征提取器作为全局专家,用于通用特征提取。局部门控网络为每个数据样本上两个专家提取的表示生成个性化权重。这三个模型构成一个局部异构MoE。加权混合表示融合了通用特征和个性化特征,并由携带个性化预测信息的局部异构大模型头处理。MoE和预测头同步更新。其次,训练后的局部同构小型特征提取器被发送到服务器,通过聚合实现跨客户端信息融合。总体而言,pFedMoE在细粒度数据级别增强了局部模型的个性化能力,同时支持模型异构性。