With evolving data regulations, machine unlearning (MU) has become an important tool for fostering trust and safety in today's AI models. However, existing MU methods focusing on data and/or weight perspectives often grapple with limitations in unlearning accuracy, stability, and cross-domain applicability. To address these challenges, we introduce the concept of 'weight saliency' in MU, drawing parallels with input saliency in model explanation. This innovation directs MU's attention toward specific model weights rather than the entire model, improving effectiveness and efficiency. The resultant method that we call saliency unlearning (SalUn) narrows the performance gap with 'exact' unlearning (model retraining from scratch after removing the forgetting dataset). To the best of our knowledge, SalUn is the first principled MU approach adaptable enough to effectively erase the influence of forgetting data, classes, or concepts in both image classification and generation. For example, SalUn yields a stability advantage in high-variance random data forgetting, e.g., with a 0.2% gap compared to exact unlearning on the CIFAR-10 dataset. Moreover, in preventing conditional diffusion models from generating harmful images, SalUn achieves nearly 100% unlearning accuracy, outperforming current state-of-the-art baselines like Erased Stable Diffusion and Forget-Me-Not.
翻译:随着数据法规的不断演进,机器遗忘(MU)已成为促进当代AI模型可信度与安全性的重要工具。然而,现有聚焦于数据和/或权重视角的MU方法,在遗忘精度、稳定性及跨领域适用性方面常面临局限。为应对这些挑战,我们提出MU中“权重显著性”的概念,类比模型解释中的输入显著性。这一创新将MU的注意力导向特定模型权重而非整个模型,从而提升效果与效率。由此产生的方法——显著性遗忘(SalUn),缩小了与“精确”遗忘(移除遗忘数据集后从头开始重训练模型)的性能差距。据我们所知,SalUn是首个具有足够适应性的原则性MU方法,能在图像分类与生成任务中有效消除遗忘数据、类别或概念的影响。例如,在高方差随机数据遗忘场景中,SalUn展现出稳定性优势——在CIFAR-10数据集上,其与精确遗忘的精度差距仅为0.2%。此外,在阻止条件扩散模型生成有害图像的任务中,SalUn实现了近100%的遗忘准确率,优于当前最先进的基线方法(如Erased Stable Diffusion与Forget-Me-Not)。