Safeguarding data from unauthorized exploitation is vital for privacy and security, especially in recent rampant research in security breach such as adversarial/membership attacks. To this end, \textit{unlearnable examples} (UEs) have been recently proposed as a compelling protection, by adding imperceptible perturbation to data so that models trained on them cannot classify them accurately on original clean distribution. Unfortunately, we find UEs provide a false sense of security, because they cannot stop unauthorized users from utilizing other unprotected data to remove the protection, by turning unlearnable data into learnable again. Motivated by this observation, we formally define a new threat by introducing \textit{learnable unauthorized examples} (LEs) which are UEs with their protection removed. The core of this approach is a novel purification process that projects UEs onto the manifold of LEs. This is realized by a new joint-conditional diffusion model which denoises UEs conditioned on the pixel and perceptual similarity between UEs and LEs. Extensive experiments demonstrate that LE delivers state-of-the-art countering performance against both supervised UEs and unsupervised UEs in various scenarios, which is the first generalizable countermeasure to UEs across supervised learning and unsupervised learning.
翻译:摘要:保护数据免受未经授权的利用对于隐私和安全至关重要,尤其是在近期层出不穷的安全漏洞研究(如对抗性攻击/成员推断攻击)中。为此,研究者近期提出了一种名为“不可学习样本”的强力保护方法,通过向数据中添加人眼不可察觉的扰动,使得基于这些数据训练的模型无法准确识别原始干净分布中的样本。然而,我们发现不可学习样本提供了一种虚假的安全感,因为它们无法阻止未授权用户利用其他未受保护的数据来移除这种保护,从而将不可学习数据重新转化为可学习数据。受此观察启发,我们形式化定义了一种新威胁——引入“可学习未授权样本”,即被移除保护后的不可学习样本。该方法的核心是一种新型纯化过程,可将不可学习样本投影到可学习未授权样本的流形上。这通过一种新的联合条件扩散模型实现,该模型基于不可学习样本与可学习未授权样本之间的像素相似性与感知相似性,对不可学习样本进行去噪处理。大量实验表明,可学习未授权样本在各种场景下对监督式不可学习样本和无监督式不可学习样本均达到了最先进的对抗性能,这是首个在监督学习与无监督学习领域均可泛化应用的不可学习样本对抗方案。