The Noisy-SGD algorithm is widely used for privately training machine learning models. Traditional privacy analyses of this algorithm assume that the internal state is publicly revealed, resulting in privacy loss bounds that increase indefinitely with the number of iterations. However, recent findings have shown that if the internal state remains hidden, then the privacy loss might remain bounded. Nevertheless, this remarkable result heavily relies on the assumption of (strong) convexity of the loss function. It remains an important open problem to further relax this condition while proving similar convergent upper bounds on the privacy loss. In this work, we address this problem for DP-SGD, a popular variant of Noisy-SGD that incorporates gradient clipping to limit the impact of individual samples on the training process. Our findings demonstrate that the privacy loss of projected DP-SGD converges exponentially fast, without requiring convexity or smoothness assumptions on the loss function. In addition, we analyze the privacy loss of regularized (unprojected) DP-SGD. To obtain these results, we directly analyze the hockey-stick divergence between coupled stochastic processes by relying on non-linear data processing inequalities.
翻译:Noisy-SGD算法被广泛用于私有机器学习模型训练。传统上对该算法的隐私分析假设内部状态公开暴露,导致隐私损失界限随迭代次数无限增长。然而,近期研究表明,若内部状态保持隐藏,隐私损失可能保持有界。但这一重要成果高度依赖于损失函数的(强)凸性假设。如何在放宽该条件的同时证明类似的隐私损失收敛上界,仍是一个重要的开放问题。本文针对DP-SGD(Noisy-SGD的流行变体,通过梯度裁剪限制单个样本对训练过程的影响)解决了该问题。研究发现,投影DP-SGD的隐私损失呈指数级快速收敛,且无需对损失函数做凸性或光滑性假设。此外,我们还分析了正则化(非投影)DP-SGD的隐私损失。为得到这些结果,我们通过利用非线性数据处理不等式,直接分析了耦合随机过程间的曲棍球stick散度。