Federated multi-view clustering offers the potential to develop a global clustering model using data distributed across multiple devices. However, current methods face challenges due to the absence of label information and the paramount importance of data privacy. A significant issue is the feature heterogeneity across multi-view data, which complicates the effective mining of complementary clustering information. Additionally, the inherent incompleteness of multi-view data in a distributed setting can further complicate the clustering process. To address these challenges, we introduce a federated incomplete multi-view clustering framework with heterogeneous graph neural networks (FIM-GNNs). In the proposed FIM-GNNs, autoencoders built on heterogeneous graph neural network models are employed for feature extraction of multi-view data at each client site. At the server level, heterogeneous features from overlapping samples of each client are aggregated into a global feature representation. Global pseudo-labels are generated at the server to enhance the handling of incomplete view data, where these labels serve as a guide for integrating and refining the clustering process across different data views. Comprehensive experiments have been conducted on public benchmark datasets to verify the performance of the proposed FIM-GNNs in comparison with state-of-the-art algorithms.
翻译:联邦多视图聚类具备利用多设备分布式数据构建全局聚类模型的潜力。然而,现有方法因缺乏标签信息及数据隐私的极端重要性而面临挑战。一个关键问题在于多视图数据的特征异质性,这增加了有效挖掘互补聚类信息的难度。此外,分布式环境下多视图数据固有的不完全性可能进一步加剧聚类过程的复杂性。为应对这些挑战,本文提出一种基于异构图神经网络的联邦不完全多视图聚类框架(FIM-GNNs)。在所提出的FIM-GNNs中,各客户端站点采用基于异构图神经网络模型构建的自编码器进行多视图数据的特征提取。在服务器端,各客户端重叠样本的异构特征被聚合为全局特征表示。服务器端生成的全局伪标签用于增强对不完全视图数据的处理能力,这些标签将作为整合与优化跨视图聚类过程的指导依据。通过在公开基准数据集上的系统实验,验证了所提FIM-GNNs相较于前沿算法的性能优势。