In all state-of-the-art sketching and coreset techniques for clustering, as well as in the best known fixed-parameter tractable approximation algorithms, randomness plays a key role. For the classic $k$-median and $k$-means problems, there are no known deterministic dimensionality reduction procedure or coreset construction that avoid an exponential dependency on the input dimension $d$, the precision parameter $\varepsilon^{-1}$ or $k$. Furthermore, there is no coreset construction that succeeds with probability $1-1/n$ and whose size does not depend on the number of input points, $n$. This has led researchers in the area to ask what is the power of randomness for clustering sketches [Feldman, WIREs Data Mining Knowl. Discov'20]. Similarly, the best approximation ratio achievable deterministically without a complexity exponential in the dimension are $\Omega(1)$ for both $k$-median and $k$-means, even when allowing a complexity FPT in the number of clusters $k$. This stands in sharp contrast with the $(1+\varepsilon)$-approximation achievable in that case, when allowing randomization. In this paper, we provide deterministic sketches constructions for clustering, whose size bounds are close to the best-known randomized ones. We also construct a deterministic algorithm for computing $(1+\varepsilon)$-approximation to $k$-median and $k$-means in high dimensional Euclidean spaces in time $2^{k^2/\varepsilon^{O(1)}} poly(nd)$, close to the best randomized complexity. Furthermore, our new insights on sketches also yield a randomized coreset construction that uses uniform sampling, that immediately improves over the recent results of [Braverman et al. FOCS '22] by a factor $k$.
翻译:在所有最先进的聚类概要图与核心集技术中,以及已知的最佳固定参数可追踪近似算法中,随机性均扮演着关键角色。对于经典的$k$-中位数和$k$-均值问题,目前尚无已知的确定性降维过程或核心集构造能避免对输入维度$d$、精度参数$\varepsilon^{-1}$或$k$的指数依赖。此外,尚无一种核心集构造能以概率$1-1/n$成功且其大小不依赖于输入点数$n$。这促使该领域研究者探究随机性在聚类概要图中的能力[Feldman, WIREs Data Mining Knowl. Discov'20]。类似地,在不依赖维度指数复杂度的确定性条件下,$k$-中位数和$k$-均值可达到的最佳近似比均为$\Omega(1)$,即使允许复杂度关于聚类数$k$为FPT。这与允许随机化时在该情形下可达到的$(1+\varepsilon)$-近似形成鲜明对比。本文为聚类问题提供了确定性概要图构造,其大小上界接近已知最佳随机构造。我们还构建了一种确定性算法,可在$2^{k^2/\varepsilon^{O(1)}} poly(nd)$时间内计算高维欧氏空间中$k$-中位数和$k$-均值的$(1+\varepsilon)$-近似,该复杂度接近最佳随机复杂度。此外,我们对概要图的新见解还催生了一种使用均匀采样的随机化核心集构造,该构造立即将[Braverman等人, FOCS '22]的最新结果改进了因子$k$。