Recent efforts have been made to integrate self-supervised learning (SSL) with the framework of federated learning (FL). One unique challenge of federated self-supervised learning (FedSSL) is that the global objective of FedSSL usually does not equal the weighted sum of local SSL objectives. Consequently, conventional approaches, such as federated averaging (FedAvg), fail to precisely minimize the FedSSL global objective, often resulting in suboptimal performance, especially when data is non-i.i.d.. To fill this gap, we propose a provable FedSSL algorithm, named FedSC, based on the spectral contrastive objective. In FedSC, clients share correlation matrices of data representations in addition to model weights periodically, which enables inter-client contrast of data samples in addition to intra-client contrast and contraction, resulting in improved quality of data representations. Differential privacy (DP) protection is deployed to control the additional privacy leakage on local datasets when correlation matrices are shared. We also provide theoretical analysis on the convergence and extra privacy leakage. The experimental results validate the effectiveness of our proposed algorithm.
翻译:近年来,自监督学习与联邦学习框架的融合研究已取得进展。联邦自监督学习面临的一个独特挑战在于,其全局目标通常不等于各客户端本地SSL目标的加权和。因此,联邦平均等传统方法无法精确最小化FedSSL全局目标,尤其在数据非独立同分布时常导致性能次优。为弥补这一不足,我们提出基于谱对比目标的可证明FedSSL算法FedSC。该算法中,客户端除了周期性共享模型权重外,还共享数据表示的协相关矩阵,这不仅实现客户端内对比与收缩,还能进行跨客户端数据样本对比,从而提升数据表示质量。通过部署差分隐私保护机制,可控制在共享协相关矩阵时对本地数据集产生的额外隐私泄露风险。我们进一步提供了算法收敛性与额外隐私泄露的理论分析,实验结果验证了所提算法的有效性。