While generation of synthetic data under differential privacy (DP) has received a lot of attention in the data privacy community, analysis of synthetic data has received much less. Existing work has shown that simply analysing DP synthetic data as if it were real does not produce valid inferences of population-level quantities. For example, confidence intervals become too narrow, which we demonstrate with a simple experiment. We tackle this problem by combining synthetic data analysis techniques from the field of multiple imputation (MI), and synthetic data generation using noise-aware (NA) Bayesian modeling into a pipeline NA+MI that allows computing accurate uncertainty estimates for population-level quantities from DP synthetic data. To implement NA+MI for discrete data generation using the values of marginal queries, we develop a novel noise-aware synthetic data generation algorithm NAPSU-MQ using the principle of maximum entropy. Our experiments demonstrate that the pipeline is able to produce accurate confidence intervals from DP synthetic data. The intervals become wider with tighter privacy to accurately capture the additional uncertainty stemming from DP noise.
翻译:差分隐私(DP)下的合成数据生成在数据隐私领域受到了广泛关注,但合成数据的分析却较少被研究。已有工作表明,简单地将DP合成数据当作真实数据进行分析,无法对总体层面的参数进行有效推断。例如,置信区间会变得过窄,我们通过一个简单实验验证了这一点。为解决此问题,我们将多重插补(MI)领域的合成数据分析技术与基于噪声感知(NA)贝叶斯建模的合成数据生成方法相结合,构建了名为NA+MI的流水线,该流水线能够从DP合成数据中计算出总体层面参数的精确不确定性估计。为了利用边际查询的值实现离散数据的生成,我们基于最大熵原理开发了一种新型噪声感知合成数据生成算法NAPSU-MQ。实验表明,该流水线能够从DP合成数据中生成准确的置信区间。随着隐私保护强度的增加,区间会相应变宽,以准确反映由DP噪声带来的额外不确定性。