We consider the problem of spectral clustering under group fairness constraints, where samples from each sensitive group are approximately proportionally represented in each cluster. Traditional fair spectral clustering (FSC) methods consist of two consecutive stages, i.e., performing fair spectral embedding on a given graph and conducting $k$means to obtain discrete cluster labels. However, in practice, the graph is usually unknown, and we need to construct the underlying graph from potentially noisy data, the quality of which inevitably affects subsequent fair clustering performance. Furthermore, performing FSC through separate steps breaks the connections among these steps, leading to suboptimal results. To this end, we first theoretically analyze the effect of the constructed graph on FSC. Motivated by the analysis, we propose a novel graph construction method with a node-adaptive graph filter to learn graphs from noisy data. Then, all independent stages of conventional FSC are integrated into a single objective function, forming an end-to-end framework that inputs raw data and outputs discrete cluster labels. An algorithm is developed to jointly and alternately update the variables in each stage. Finally, we conduct extensive experiments on synthetic, benchmark, and real data, which show that our model is superior to state-of-the-art fair clustering methods.
翻译:我们研究在群体公平约束下的谱聚类问题,其中每个敏感群体中的样本在每个簇中大致按比例代表。传统的公平谱聚类方法由两个连续阶段组成,即在给定图上执行公平谱嵌入,然后进行$k$均值聚类以获得离散的簇标签。然而,在实际中,图通常是未知的,我们需要从可能带有噪声的数据中构建底层图,其质量不可避免地影响后续的公平聚类性能。此外,通过分离步骤执行公平谱聚类会切断这些步骤之间的联系,导致次优结果。为此,我们首先从理论上分析了构建图对公平谱聚类的影响。受此分析启发,我们提出了一种新的图构建方法,该方法采用节点自适应图滤波器从噪声数据中学习图。然后,将传统公平谱聚类的所有独立阶段整合到一个单一目标函数中,形成一个端到端框架,该框架输入原始数据并输出离散簇标签。我们开发了一种算法来联合且交替地更新每个阶段的变量。最后,我们在合成数据、基准数据和真实数据上进行了大量实验,结果表明我们的模型优于最先进的公平聚类方法。