Model stealing attacks, where adversaries create high-fidelity surrogate models, are a significant threat to the intellectual property of machine learning services. Conventional wisdom suggests these surrogates could provide adversaries with economic leverage comparable to the original service providers. This paper challenges this assumption by evaluating model stealing attacks beyond mere fidelity to the target model. Because query-based extraction provides only partial supervision of the target's input-output behavior, the surrogate is not uniquely identified: many near-optimal surrogates can achieve comparable fidelity while differing in deployment-relevant properties. Instead of performing a classic learning-based model stealing attack, we compute the Rashomon Set (i.e., the set of almost-equally-accurate models) of surrogate models, and evaluate its diversity using multiplicity metrics (ambiguity, discrepancy, and Rashomon Capacity) and group fairness metrics. Across tabular, medical imaging, and NLP tasks, our experiments on real-world datasets reveal that despite exhibiting similar fidelity to the target model, surrogate models can display significant variances in other critical performance metrics. These findings cast doubt on the presumed equivalence between high-fidelity surrogates and the target model in practical deployment scenarios.


翻译:模型窃取攻击(即攻击者构建高保真替代模型的行为)对机器学习服务的知识产权构成重大威胁。传统观点认为,这些替代模型可使攻击者获得与原始服务提供者相当的经济优势。本文通过评估超越目标模型保真度的模型窃取攻击,对这一假设提出质疑。由于基于查询的提取仅提供目标输入输出行为的部分监督,替代模型并非唯一确定:大量近优替代模型能在保持相似保真度的同时,展现出部署相关属性的显著差异。我们不采用经典的学习型模型窃取方法,而是计算替代模型的Rashomon集合(即精度几乎相等的模型集合),并利用多样性指标(模糊性、差异性与Rashomon容量)及群体公平性指标评估其多样性。在表格数据、医学影像及自然语言处理任务中,基于真实数据集的实验表明,尽管替代模型与目标模型在保真度上表现相似,但在其他关键性能指标上可能存在显著差异。这些发现对高保真替代模型在实际部署场景中与目标模型等价性的传统认知提出了质疑。

0
下载
关闭预览

相关内容

深度学习模型反演攻击与防御:全面综述
专知会员服务
27+阅读 · 2025年2月3日
深度学习模型安全:威胁与防御,176页pdf
专知会员服务
28+阅读 · 2024年12月13日
多视角看大模型安全及实践
专知会员服务
70+阅读 · 2024年4月1日
专知会员服务
24+阅读 · 2021年8月22日
专知会员服务
49+阅读 · 2021年5月17日
专知会员服务
57+阅读 · 2020年12月28日
机器学习模型安全与隐私研究综述
专知会员服务
116+阅读 · 2020年11月12日
模型攻击:鲁棒性联邦学习研究的最新进展
机器之心
35+阅读 · 2020年6月3日
一文读懂机器学习模型的选择与取舍
DBAplus社群
13+阅读 · 2019年8月25日
用模型不确定性理解模型
论智
11+阅读 · 2018年9月5日
【学界】机器学习模型的“可解释性”到底有多重要?
GAN生成式对抗网络
12+阅读 · 2018年3月3日
机器学习模型的“可解释性”到底有多重要?
中国科学院自动化研究所
20+阅读 · 2018年3月1日
推荐|机器学习中的模型评价、模型选择和算法选择!
全球人工智能
10+阅读 · 2018年2月5日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
深度学习模型反演攻击与防御:全面综述
专知会员服务
27+阅读 · 2025年2月3日
深度学习模型安全:威胁与防御,176页pdf
专知会员服务
28+阅读 · 2024年12月13日
多视角看大模型安全及实践
专知会员服务
70+阅读 · 2024年4月1日
专知会员服务
24+阅读 · 2021年8月22日
专知会员服务
49+阅读 · 2021年5月17日
专知会员服务
57+阅读 · 2020年12月28日
机器学习模型安全与隐私研究综述
专知会员服务
116+阅读 · 2020年11月12日
相关基金
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员