Model inversion attacks involve reconstructing the training data of a target model, which raises serious privacy concerns for machine learning models. However, these attacks, especially learning-based methods, are likely to suffer from low attack accuracy, i.e., low classification accuracy of these reconstructed data by machine learning classifiers. Recent studies showed an alternative strategy of model inversion attacks, GAN-based optimization, can improve the attack accuracy effectively. However, these series of GAN-based attacks reconstruct only class-representative training data for a class, whereas learning-based attacks can reconstruct diverse data for different training data in each class. Hence, in this paper, we propose a new training paradigm for a learning-based model inversion attack that can achieve higher attack accuracy in a black-box setting. First, we regularize the training process of the attack model with an added semantic loss function and, second, we inject adversarial examples into the training data to increase the diversity of the class-related parts (i.e., the essential features for classification tasks) in training data. This scheme guides the attack model to pay more attention to the class-related parts of the original data during the data reconstruction process. The experimental results show that our method greatly boosts the performance of existing learning-based model inversion attacks. Even when no extra queries to the target model are allowed, the approach can still improve the attack accuracy of reconstructed data. This new attack shows that the severity of the threat from learning-based model inversion adversaries is underestimated and more robust defenses are required.
翻译:模型反转攻击涉及重建目标模型的训练数据,这引发了机器学习模型的严重隐私问题。然而,这些攻击(尤其是基于学习的方法)往往面临攻击精度低的问题,即机器学习分类器对这些重建数据的分类准确率较低。近期研究表明,基于生成对抗网络优化的模型反转攻击替代策略可有效提升攻击精度。但这类基于GAN的攻击只能重建某个类别的类代表性训练数据,而基于学习的攻击能够重建每个类别中不同训练样本的多样化数据。为此,本文提出一种新的基于学习的模型反转攻击训练范式,可在黑盒设置下实现更高攻击精度。首先,我们通过添加语义损失函数对攻击模型的训练过程进行正则化;其次,向训练数据中注入对抗样本以增加训练数据中类别相关部分(即分类任务的关键特征)的多样性。该方案引导攻击模型在数据重建过程中更关注原始数据的类别相关部分。实验结果表明,我们的方法显著提升了现有基于学习的模型反转攻击性能。即使在禁止向目标模型发起额外查询的情况下,该方法仍能提高重建数据的攻击精度。这种新型攻击表明,基于学习的模型反转攻击者威胁的严重性被低估,亟需更鲁棒的防御措施。