Differential privacy with gradual expiration models the setting where data items arrive in a stream and at a given time $t$ the privacy loss guaranteed for a data item seen at time $(t-d)$ is $\epsilon g(d)$, where $g$ is a monotonically non-decreasing function. We study the fundamental $\textit{continual (binary) counting}$ problem where each data item consists of a bit, and the algorithm needs to output at each time step the sum of all the bits streamed so far. For a stream of length $T$ and privacy $\textit{without}$ expiration continual counting is possible with maximum (over all time steps) additive error $O(\log^2(T)/\varepsilon)$ and the best known lower bound is $\Omega(\log(T)/\varepsilon)$; closing this gap is a challenging open problem. We show that the situation is very different for privacy with gradual expiration by giving upper and lower bounds for a large set of expiration functions $g$. Specifically, our algorithm achieves an additive error of $ O(\log(T)/\epsilon)$ for a large set of privacy expiration functions. We also give a lower bound that shows that if $C$ is the additive error of any $\epsilon$-DP algorithm for this problem, then the product of $C$ and the privacy expiration function after $2C$ steps must be $\Omega(\log(T)/\epsilon)$. Our algorithm matches this lower bound as its additive error is $O(\log(T)/\epsilon)$, even when $g(2C) = O(1)$. Our empirical evaluation shows that we achieve a slowly growing privacy loss with significantly smaller empirical privacy loss for large values of $d$ than a natural baseline algorithm.
翻译:具有逐步过期模型的差分隐私描述了这样一种场景:数据项以流形式到达,在给定时间 $t$ 时,对于在时间 $(t-d)$ 出现的数据项,其隐私损失保证为 $\epsilon g(d)$,其中 $g$ 是单调非递减函数。我们研究基础的$\textit{持续(二进制)计数}$问题,其中每个数据项由一个比特组成,算法需要在每个时间步输出到目前为止所有流式比特的总和。对于长度为 $T$ 且$\textit{无}$过期的隐私流,持续计数的最大(所有时间步上)加性误差为 $O(\log^2(T)/\varepsilon)$,而已知最佳下界为 $\Omega(\log(T)/\varepsilon)$;缩小这一差距是一个具有挑战性的开放问题。我们证明,对于具有逐步过期模型的隐私,情况截然不同:我们针对一大类过期函数 $g$ 给出了上下界。具体而言,我们的算法对于一大类隐私过期函数实现了 $O(\log(T)/\epsilon)$ 的加性误差。我们还给出了一个下界,表明如果 $C$ 是任何 $\epsilon$-DP 算法针对此问题的加性误差,那么在 $2C$ 步后,$C$ 与隐私过期函数的乘积必须为 $\Omega(\log(T)/\epsilon)$。我们的算法匹配了这一下界,因为即使当 $g(2C) = O(1)$ 时,其加性误差也为 $O(\log(T)/\epsilon)$。我们的实证评估表明,与自然基线算法相比,我们在 $d$ 较大时实现了显著更小的经验隐私损失,且隐私损失增长缓慢。