When training a machine learning model with differential privacy, one sets a privacy budget. This budget represents a maximal privacy violation that any user is willing to face by contributing their data to the training set. We argue that this approach is limited because different users may have different privacy expectations. Thus, setting a uniform privacy budget across all points may be overly conservative for some users or, conversely, not sufficiently protective for others. In this paper, we capture these preferences through individualized privacy budgets. To demonstrate their practicality, we introduce a variant of Differentially Private Stochastic Gradient Descent (DP-SGD) which supports such individualized budgets. DP-SGD is the canonical approach to training models with differential privacy. We modify its data sampling and gradient noising mechanisms to arrive at our approach, which we call Individualized DP-SGD (IDP-SGD). Because IDP-SGD provides privacy guarantees tailored to the preferences of individual users and their data points, we find it empirically improves privacy-utility trade-offs.
翻译:在采用差分隐私技术训练机器学习模型时,通常需要设定一个隐私预算。该预算代表任何用户因贡献数据至训练集而愿意承受的最大隐私泄露风险。我们认为这种设定存在局限性,因为不同用户可能具有差异化的隐私期望。因此,对所有数据点采用统一的隐私预算,可能会导致对部分用户过度保守,而对其余用户保护不足。本文通过引入个性化隐私预算来捕捉这些偏好差异。为验证其可行性,我们提出一种支持个性化预算的差分隐私随机梯度下降(DP-SGD)变体。作为差分隐私模型训练的经典方法,DP-SGD 的数据采样与梯度加噪机制被我们改进为"个性化DP-SGD"(IDP-SGD)。由于IDP-SGD能为不同用户及其数据点提供符合其偏好设定的隐私保障,实验证明该方法有效提升了隐私-效用平衡的性能。