Clustering is a fundamental building block of modern statistical analysis pipelines. Fair clustering has seen much attention from the machine learning community in recent years. We are some of the first to study fairness in the context of hierarchical clustering, after the results of Ahmadian et al. from NeurIPS in 2020. We evaluate our results using Dasgupta's cost function, perhaps one of the most prevalent theoretical metrics for hierarchical clustering evaluation. Our work vastly improves the previous $O(n^{5/6}poly\log(n))$ fair approximation for cost to a near polylogarithmic $O(n^\delta poly\log(n))$ fair approximation for any constant $\delta\in(0,1)$. This result establishes a cost-fairness tradeoff and extends to broader fairness constraints than the previous work. We also show how to alter existing hierarchical clusterings to guarantee fairness and cluster balance across any level in the hierarchy.
翻译:聚类是现代统计分析流程中的基本构建模块。近年来,公平聚类引起了机器学习社区的广泛关注。在Ahmadian等人2020年发表于NeurIPS的研究成果之后,我们率先探讨了层次聚类中的公平性问题。我们使用Dasgupta代价函数评估结果,该函数或许是层次聚类评估中最主流的理论指标之一。我们的工作将先前代价的$O(n^{5/6}\mathrm{poly}\log(n))$公平近似大幅提升至接近多对数级别的$O(n^{\delta}\mathrm{poly}\log(n))$公平近似,其中$\delta\in(0,1)$为任意常数。这一结果建立了代价-公平性权衡,并相较于先前工作扩展到更广泛的公平约束。我们还展示了如何修改现有层次聚类,以在层次结构的任意层级保证公平性与聚类平衡性。