We formalize the problem of machine unlearning as design of efficient unlearning algorithms corresponding to learning algorithms which perform a selection of adaptive queries from structured query classes. We give efficient unlearning algorithms for linear and prefix-sum query classes. As applications, we show that unlearning in many problems, in particular, stochastic convex optimization (SCO), can be reduced to the above, yielding improved guarantees for the problem. In particular, for smooth Lipschitz losses and any $\rho>0$, our results yield an unlearning algorithm with excess population risk of $\tilde O\big(\frac{1}{\sqrt{n}}+\frac{\sqrt{d}}{n\rho}\big)$ with unlearning query (gradient) complexity $\tilde O(\rho \cdot \text{Retraining Complexity})$, where $d$ is the model dimensionality and $n$ is the initial number of samples. For non-smooth Lipschitz losses, we give an unlearning algorithm with excess population risk $\tilde O\big(\frac{1}{\sqrt{n}}+\big(\frac{\sqrt{d}}{n\rho}\big)^{1/2}\big)$ with the same unlearning query (gradient) complexity. Furthermore, in the special case of Generalized Linear Models (GLMs), such as those in linear and logistic regression, we get dimension-independent rates of $\tilde O\big(\frac{1}{\sqrt{n}} +\frac{1}{(n\rho)^{2/3}}\big)$ and $\tilde O\big(\frac{1}{\sqrt{n}} +\frac{1}{(n\rho)^{1/3}}\big)$ for smooth Lipschitz and non-smooth Lipschitz losses respectively. Finally, we give generalizations of the above from one unlearning request to \textit{dynamic} streams consisting of insertions and deletions.
翻译:我们将机器遗忘问题形式化为:针对从结构化查询类中选择自适应查询的学习算法,设计高效的遗忘算法。我们为线性查询类和前缀和查询类给出了高效的遗忘算法。作为应用,我们展示了多种问题中的遗忘过程(尤其是随机凸优化)可归约至上述框架,从而为该问题带来更优的保证。具体而言,对于光滑Lipschitz损失和任意$\rho>0$,我们的结果给出了一个遗忘算法,其超额总体风险为$\tilde O\big(\frac{1}{\sqrt{n}}+\frac{\sqrt{d}}{n\rho}\big)$,且遗忘查询(梯度)复杂度为$\tilde O(\rho \cdot \text{重训练复杂度})$,其中$d$是模型维度,$n$是初始样本数。对于非光滑Lipschitz损失,我们给出一个遗忘算法,超额总体风险为$\tilde O\big(\frac{1}{\sqrt{n}}+\big(\frac{\sqrt{d}}{n\rho}\big)^{1/2}\big)$,且具有相同的遗忘查询(梯度)复杂度。此外,在广义线性模型的特殊情形中(如线性回归和逻辑回归),对于光滑Lipschitz和非光滑Lipschitz损失,我们分别获得了与维度无关的速率$\tilde O\big(\frac{1}{\sqrt{n}} +\frac{1}{(n\rho)^{2/3}}\big)$和$\tilde O\big(\frac{1}{\sqrt{n}} +\frac{1}{(n\rho)^{1/3}}\big)$。最后,我们将上述结果从单次遗忘请求推广至包含插入和删除操作的动态流场景。