The advent of online genomic data-sharing services has sought to enhance the accessibility of large genomic datasets by allowing queries about genetic variants, such as summary statistics, aiding care providers in distinguishing between spurious genomic variations and those with clinical significance. However, numerous studies have demonstrated that even sharing summary genomic information exposes individual members of such datasets to a significant privacy risk due to membership inference attacks. While several approaches have emerged that reduce privacy risks by adding noise or reducing the amount of information shared, these typically assume non-adaptive attacks that use likelihood ratio test (LRT) statistics. We propose a Bayesian game-theoretic framework for optimal privacy-utility tradeoff in the sharing of genomic summary statistics. Our first contribution is to prove that a very general Bayesian attacker model that anchors our game-theoretic approach is more powerful than the conventional LRT-based threat models in that it induces worse privacy loss for the defender who is modeled as a von Neumann-Morgenstern (vNM) decision-maker. We show this to be true even when the attacker uses a non-informative subjective prior. Next, we present an analytically tractable approach to compare the Bayesian attacks with arbitrary subjective priors and the Neyman-Pearson optimal LRT attacks under the Gaussian mechanism common in differential privacy frameworks. Finally, we propose an approach for approximating Bayes-Nash equilibria of the game using deep neural network generators to implicitly represent player mixed strategies. Our experiments demonstrate that the proposed game-theoretic framework yields both stronger attacks and stronger defense strategies than the state of the art.
翻译:在线基因组数据共享服务的出现旨在通过允许查询遗传变异(如摘要统计量)来增强大型基因组数据集的可访问性,帮助医疗服务提供者区分无意义的基因组变异与具有临床意义的变异。然而,多项研究表明,由于成员推断攻击的存在,即使共享基因组摘要信息也会使此类数据集的个体成员面临重大的隐私风险。尽管已出现多种通过添加噪声或减少共享信息量来降低隐私风险的方法,但这些方法通常假设攻击者使用似然比检验统计量进行非自适应攻击。我们提出了一个贝叶斯博弈论框架,用于在基因组摘要统计共享中实现最优的隐私-效用权衡。我们的第一个贡献是证明:作为我们博弈论方法基础的非常通用的贝叶斯攻击者模型,比传统的基于似然比检验的威胁模型更强大,因为它会对被建模为冯·诺依曼-摩根斯坦决策者的防御方造成更严重的隐私损失。我们证明即使攻击者使用无信息的主观先验,这一结论依然成立。其次,我们提出了一种解析可处理的方法,用于在差分隐私框架中常见的高斯机制下,比较具有任意主观先验的贝叶斯攻击与奈曼-皮尔逊最优似然比检验攻击。最后,我们提出了一种利用深度神经网络生成器隐式表示参与者混合策略来逼近博弈贝叶斯-纳什均衡的方法。实验表明,所提出的博弈论框架能够产生比现有技术更强大的攻击策略和防御策略。