This paper studies the fair range clustering problem in which the data points are from different demographic groups and the goal is to pick $k$ centers with the minimum clustering cost such that each group is at least minimally represented in the centers set and no group dominates the centers set. More precisely, given a set of $n$ points in a metric space $(P,d)$ where each point belongs to one of the $\ell$ different demographics (i.e., $P = P_1 \uplus P_2 \uplus \cdots \uplus P_\ell$) and a set of $\ell$ intervals $[\alpha_1, \beta_1], \cdots, [\alpha_\ell, \beta_\ell]$ on desired number of centers from each group, the goal is to pick a set of $k$ centers $C$ with minimum $\ell_p$-clustering cost (i.e., $(\sum_{v\in P} d(v,C)^p)^{1/p}$) such that for each group $i\in \ell$, $|C\cap P_i| \in [\alpha_i, \beta_i]$. In particular, the fair range $\ell_p$-clustering captures fair range $k$-center, $k$-median and $k$-means as its special cases. In this work, we provide an $O(1)$-approximation algorithm for the fair range $\ell_p$-clustering that picks at most $k+2\ell$ centers and may only violate the upper bound of each demographic group by at most an additive term of $2$.
翻译:本文研究了公平范围聚类问题,其中数据点来自不同的人口统计组,目标是在最小化聚类成本的前提下选取$k$个中心,使得每个组在中心集合中至少获得最小代表性,且没有组主导中心集合。更精确地说,给定度量空间$(P,d)$中的一组$n$个点,每个点属于$\ell$个不同人口统计组之一(即$P = P_1 \uplus P_2 \uplus \cdots \uplus P_\ell$),以及一组关于每组期望中心数量的$\ell$个区间$[\alpha_1, \beta_1], \cdots, [\alpha_\ell, \beta_\ell]$,目标是最小化$\ell_p$聚类成本(即$(\sum_{v\in P} d(v,C)^p)^{1/p}$)选取一组$k$个中心$C$,使得对于每个组$i\in \ell$,有$|C\cap P_i| \in [\alpha_i, \beta_i]$。特别地,公平范围$\ell_p$聚类将公平范围$k$-中心、$k$-中位数和$k$-均值作为其特例。在本文中,我们为公平范围$\ell_p$聚类提供了一个$O(1)$-近似算法,该算法选取至多$k+2\ell$个中心,并且对每个人口统计组的上界违反程度至多为$2$的加法项。