The canonical algorithm for differentially private mean estimation is to first clip the samples to a bounded range and then add noise to their empirical mean. Clipping controls the sensitivity and, hence, the variance of the noise that we add for privacy. But clipping also introduces statistical bias. We prove that this tradeoff is inherent: no algorithm can simultaneously have low bias, low variance, and low privacy loss for arbitrary distributions. On the positive side, we show that unbiased mean estimation is possible under approximate differential privacy if we assume that the distribution is symmetric. Furthermore, we show that, even if we assume that the data is sampled from a Gaussian, unbiased mean estimation is impossible under pure or concentrated differential privacy.
翻译:差分隐私均值估计的经典算法首先将样本裁剪至有界范围,然后向其实验均值添加噪声。裁剪控制了灵敏度,进而控制了为隐私保护而添加的噪声方差。但裁剪同时引入了统计偏差。我们证明这种权衡是固有的:对于任意分布,任何算法都无法同时实现低偏差、低方差和低隐私损失。从积极方面看,我们证明若假设分布对称,则可在近似差分隐私下实现无偏均值估计。此外,即使假设数据采样自高斯分布,纯差分隐私或集中差分隐私下的无偏均值估计仍不可能实现。