The notion of replicable algorithms was introduced in Impagliazzo et al. [STOC '22] to describe randomized algorithms that are stable under the resampling of their inputs. More precisely, a replicable algorithm gives the same output with high probability when its randomness is fixed and it is run on a new i.i.d. sample drawn from the same distribution. Using replicable algorithms for data analysis can facilitate the verification of published results by ensuring that the results of an analysis will be the same with high probability, even when that analysis is performed on a new data set. In this work, we establish new connections and separations between replicability and standard notions of algorithmic stability. In particular, we give sample-efficient algorithmic reductions between perfect generalization, approximate differential privacy, and replicability for a broad class of statistical problems. Conversely, we show any such equivalence must break down computationally: there exist statistical problems that are easy under differential privacy, but that cannot be solved replicably without breaking public-key cryptography. Furthermore, these results are tight: our reductions are statistically optimal, and we show that any computational separation between DP and replicability must imply the existence of one-way functions. Our statistical reductions give a new algorithmic framework for translating between notions of stability, which we instantiate to answer several open questions in replicability and privacy. This includes giving sample-efficient replicable algorithms for various PAC learning, distribution estimation, and distribution testing problems, algorithmic amplification of $\delta$ in approximate DP, conversions from item-level to user-level privacy, and the existence of private agnostic-to-realizable learning reductions under structured distributions.
翻译:可复制算法(replicable algorithms)的概念由Impagliazzo等人在[STOC '22]中提出,用于描述在输入数据重采样时具有稳定性的随机化算法。更精确地说,当可复制算法的随机性固定后,若从同一分布中独立同分布地抽取新样本并运行该算法,其输出结果以高概率保持一致。使用可复制算法进行数据分析,可确保分析结果(即使在新数据集上执行)以高概率保持不变,从而促进已发表结果的可验证性。本研究建立了可复制性与算法稳定性标准概念之间的新联系与区分。具体而言,对于一类广泛的统计问题,我们给出了在完美泛化、近似差分隐私与可复制性之间进行样本高效算法化归的方法。反之,我们证明此类等价性必然存在计算性断裂:存在在差分隐私下易于处理但无法在不破解公钥密码学的前提下以可复制方式解决的统计问题。此外,这些结果具有紧致性:我们的化归在统计意义上是最优的,且任何差分隐私与可复制性之间的计算分离都必然蕴含单向函数的存在性。我们的统计化归为稳定性概念之间的转换提供了新算法框架,并由此解答了可复制性与隐私领域的多个未解决问题,包括:为多种PAC学习、分布估计与分布检验问题给出样本高效的可复制算法,近似差分隐私中δ的算法放大,从条目级到用户级隐私的转换,以及结构化分布下私有不可知到可实现学习化归的存在性。