Model Inversion (MI) attacks pose a significant privacy threat by reconstructing private training data from machine learning models. While existing defenses primarily concentrate on model-centric approaches, the impact of data on MI robustness remains largely unexplored. In this work, we explore Random Erasing (RE), a technique traditionally used for improving model generalization under occlusion, and uncover its surprising effectiveness as a defense against MI attacks. Specifically, our novel feature space analysis shows that models trained with RE-images introduce a significant discrepancy between the features of MI-reconstructed images and those of the private data. At the same time, features of private images remain distinct from other classes and well-separated from different classification regions. These effects collectively degrade MI reconstruction quality and attack accuracy while maintaining reasonable natural accuracy. Furthermore, we explore two critical properties of RE including Partial Erasure and Random Location. Partial Erasure prevents the model from observing entire objects during training. We find this has a significant impact on MI, which aims to reconstruct the entire objects. Random Location of erasure plays a crucial role in achieving a strong privacy-utility trade-off. Our findings highlight RE as a simple yet effective defense mechanism that can be easily integrated with existing privacy-preserving techniques. Extensive experiments across 37 setups demonstrate that our method achieves state-of-the-art (SOTA) performance in the privacy-utility trade-off. The results consistently demonstrate the superiority of our defense over existing methods across different MI attacks, network architectures, and attack configurations. For the first time, we achieve a significant degradation in attack accuracy without a decrease in utility for some configurations.


翻译:模型反演(MI)攻击通过从机器学习模型中重建私有训练数据,构成了重大的隐私威胁。现有防御主要集中于以模型为中心的方法,而数据对MI鲁棒性的影响仍大多未被探索。在本研究中,我们探讨了随机擦除(RE)——一种传统上用于提高遮挡条件下模型泛化能力的技术——并揭示了其作为一种针对MI攻击防御手段的惊人有效性。具体而言,我们新颖的特征空间分析表明,使用RE图像训练的模型会在MI重建图像特征与私有数据特征之间引入显著差异。同时,私有图像的特征仍与其他类别保持区分性,并与不同分类区域良好分离。这些效应共同降低了MI重建质量和攻击精度,同时保持了合理的自然精度。此外,我们探索了RE的两个关键特性,包括部分擦除和随机位置。部分擦除阻止了模型在训练过程中观察到完整物体。我们发现这对旨在重建完整物体的MI产生了显著影响。擦除的随机位置在实现强隐私-效用权衡中起着关键作用。我们的研究结果凸显了RE作为一种简单而有效的防御机制,可以轻松与现有隐私保护技术集成。跨越37种实验设置的大量实验表明,我们的方法在隐私-效用权衡中达到了最先进的性能。结果一致证明了我们的防御在不同MI攻击、网络架构和攻击配置下优于现有方法。首次,我们在某些配置中实现了攻击精度的显著下降而不降低效用。

0
下载
关闭预览

相关内容

深度学习模型反演攻击与防御:全面综述
专知会员服务
27+阅读 · 2025年2月3日
预训练模型的新兴安全与隐私问题:综述与展望
专知会员服务
20+阅读 · 2024年11月13日
【CVPR2024】持续遗忘对于预训练视觉模型
专知会员服务
19+阅读 · 2024年3月20日
【CVPR2023】基于强化学习的黑盒模型反演攻击
专知会员服务
24+阅读 · 2023年4月12日
专知会员服务
24+阅读 · 2021年8月22日
专知会员服务
49+阅读 · 2021年5月17日
专知会员服务
97+阅读 · 2021年1月17日
模型攻击:鲁棒性联邦学习研究的最新进展
机器之心
35+阅读 · 2020年6月3日
您可以相信模型的不确定性吗?
TensorFlow
14+阅读 · 2020年1月31日
用模型不确定性理解模型
论智
11+阅读 · 2018年9月5日
展望:模型驱动的深度学习
人工智能学家
12+阅读 · 2018年1月23日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
7+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
14+阅读 · 8月7日
相关VIP内容
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员