Many modern statistical analysis and machine learning applications require training models on sensitive user data. Differential privacy provides a formal guarantee that individual-level information about users does not leak. In this framework, randomized algorithms inject calibrated noise into the confidential data, resulting in privacy-protected datasets or queries. However, restricting access to only the privatized data during statistical analysis makes it computationally challenging to perform valid inferences on parameters underlying the confidential data. In this work, we propose simulation-based inference methods from privacy-protected datasets. Specifically, we use neural conditional density estimators as a flexible family of distributions to approximate the posterior distribution of model parameters given the observed private query results. We illustrate our methods on discrete time-series data under an infectious disease model and on ordinary linear regression models. Illustrating the privacy-utility trade-off, our experiments and analysis demonstrate the necessity and feasibility of designing valid statistical inference procedures to correct for biases introduced by the privacy-protection mechanisms.
翻译:许多现代统计分析和机器学习应用需要在敏感用户数据上训练模型。差分隐私提供了一种形式化保证,确保用户个体层级的信息不会泄露。在该框架下,随机化算法向机密数据注入校准噪声,生成隐私保护数据集或查询结果。然而,在统计分析过程中仅能访问私有化数据,使得对机密数据背后的参数进行有效推断在计算上具有挑战性。本文提出基于隐私保护数据集的模拟推断方法。具体而言,我们将神经条件密度估计器作为一类灵活的分布族,用于近似给定观测私有查询结果时模型参数的后验分布。我们基于传染病模型下的离散时间序列数据和普通线性回归模型对所提方法进行了验证。通过展示隐私-效用权衡,我们的实验与分析表明,设计有效的统计推断过程以校正隐私保护机制引入的偏差具有必要性与可行性。