Federated clustering is an important part of the field of federated machine learning, that allows multiple data sources to collaboratively cluster their data while keeping it decentralized and preserving privacy. In this paper, we introduce a novel federated clustering algorithm, named Dynamically Weighted Federated k-means (DWF k-means), to address the challenges posed by distributed data sources and heterogeneous data. Our proposed algorithm combines the benefits of traditional clustering techniques with the privacy and scalability advantages of federated learning. It enables multiple data owners to collaboratively cluster their local data while exchanging minimal information with a central coordinator. The algorithm optimizes the clustering process by adaptively aggregating cluster assignments and centroids from each data source, thereby learning a global clustering solution that reflects the collective knowledge of the entire federated network. We conduct experiments on multiple datasets and data distribution settings to evaluate the performance of our algorithm in terms of clustering score, accuracy, and v-measure. The results demonstrate that our approach can match the performance of the centralized classical k-means baseline, and outperform existing federated clustering methods in realistic scenarios.
翻译:联邦聚类是联邦机器学习领域的重要组成部分,它允许多个数据源在保持数据去中心化并保护隐私的前提下协作完成数据聚类。本文提出一种新型联邦聚类算法——动态加权联邦k-Means(DWF k-Means),旨在解决分布式数据源和异构数据带来的挑战。该算法融合传统聚类技术的优势与联邦学习的隐私性及可扩展性,使多个数据拥有者能够在与中央协调器交换最少信息的同时,协作完成本地数据聚类。算法通过自适应聚合各数据源的聚类分配结果与质心,优化聚类过程,从而学习反映整个联邦网络集体知识的全局聚类解决方案。我们在多种数据集及数据分布场景下开展实验,以聚类评分、准确率和V-measure指标评估算法性能。结果表明,本算法能达到集中式经典k-Means基准的性能水平,并在真实场景中优于现有联邦聚类方法。