Differential Privacy (DP) is a mathematical framework that is increasingly deployed to mitigate privacy risks associated with machine learning and statistical analyses. Despite the growing adoption of DP, its technical privacy parameters do not lend themselves to an intelligible description of the real-world privacy risks associated with that deployment: the guarantee that most naturally follows from the DP definition is protection against membership inference by an adversary who knows all but one data record and has unlimited auxiliary knowledge. In many settings, this adversary is far too strong to inform how to set real-world privacy parameters. One approach for contextualizing privacy parameters is via defining and measuring the success of technical attacks, but doing so requires a systematic categorization of the relevant attack space. In this work, we offer a detailed taxonomy of attacks, showing the various dimensions of attacks and highlighting that many real-world settings have been understudied. Our taxonomy provides a roadmap for analyzing real-world deployments and developing theoretical bounds for more informative privacy attacks. We operationalize our taxonomy by using it to analyze a real-world case study, the Israeli Ministry of Health's recent release of a birth dataset using DP, showing how the taxonomy enables fine-grained threat modeling and provides insight towards making informed privacy parameter choices. Finally, we leverage the taxonomy towards defining a more realistic attack than previously considered in the literature, namely a distributional reconstruction attack: we generalize Balle et al.'s notion of reconstruction robustness to a less-informed adversary with distributional uncertainty, and extend the worst-case guarantees of DP to this average-case setting.
翻译:差分隐私(DP)是一种数学框架,正被越来越多地用于缓解机器学习和统计分析中的隐私风险。尽管DP的采用日益增长,其技术性隐私参数却难以直观描述实际部署中的真实隐私风险:DP定义最自然提供的保障,是针对知晓全部数据记录(除一条之外)且拥有无限辅助知识的攻击者实施成员推断攻击时的防护。在许多场景中,这种攻击者过于强大,无法为实际隐私参数的设置提供参考。一种将隐私参数具象化的方法是通过定义并度量技术攻击的成功率,但这需要对相关攻击空间进行系统性分类。本文提出了一种细粒度的攻击分类体系,展示了攻击的多个维度,并指出许多现实场景尚未得到充分研究。该分类体系为分析实际部署、推导更具信息量的隐私攻击的理论边界提供了路线图。我们通过将其应用于真实案例——以色列卫生部近期使用DP发布的人口出生数据集——来验证这一分类体系的有效性,展示其如何支持细粒度的威胁建模,并为明智选择隐私参数提供洞见。最后,我们利用该分类体系定义了一种比现有文献更贴近实际的攻击,即分布重构攻击:我们将Balle等人关于重构鲁棒性的概念推广到信息较少且存在分布不确定性的攻击者,并将DP的最坏情况保障扩展至这种平均情况场景。