In this paper, we design replicable algorithms in the context of statistical clustering under the recently introduced notion of replicability. A clustering algorithm is replicable if, with high probability, it outputs the exact same clusters after two executions with datasets drawn from the same distribution when its internal randomness is shared across the executions. We propose such algorithms for the statistical $k$-medians, statistical $k$-means, and statistical $k$-centers problems by utilizing approximation routines for their combinatorial counterparts in a black-box manner. In particular, we demonstrate a replicable $O(1)$-approximation algorithm for statistical Euclidean $k$-medians ($k$-means) with $\operatorname{poly}(d)$ sample complexity. We also describe a $O(1)$-approximation algorithm with an additional $O(1)$-additive error for statistical Euclidean $k$-centers, albeit with $\exp(d)$ sample complexity.
翻译:本文在近期提出的可复现性概念框架下,针对统计聚类问题设计了可复现算法。如果一个聚类算法在两次执行中(数据集均来自同一分布且内部随机性在两次执行间共享),能以高概率输出完全相同的聚类结果,则该算法具有可复现性。我们通过黑箱方式利用组合聚类问题的近似求解程序,针对统计$k$-中位数、统计$k$-均值和统计$k$-中心问题提出了相应算法。具体而言,我们展示了一种样本复杂度为$\operatorname{poly}(d)$的可复现$O(1)$-近似算法用于统计欧氏$k$-中位数($k$-均值)问题。同时,我们描述了一种针对统计欧氏$k$-中心问题的$O(1)$-近似算法(含额外$O(1)$加法误差),尽管其样本复杂度达到$\exp(d)$。