Deep neural network (DNN) models are valuable intellectual property of model owners, constituting a competitive advantage. Therefore, it is crucial to develop techniques to protect against model theft. Model ownership resolution (MOR) is a class of techniques that can deter model theft. A MOR scheme enables an accuser to assert an ownership claim for a suspect model by presenting evidence, such as a watermark or fingerprint, to show that the suspect model was stolen or derived from a source model owned by the accuser. Most of the existing MOR schemes prioritize robustness against malicious suspects, ensuring that the accuser will win if the suspect model is indeed a stolen model. In this paper, we show that common MOR schemes in the literature are vulnerable to a different, equally important but insufficiently explored, robustness concern: a malicious accuser. We show how malicious accusers can successfully make false claims against independent suspect models that were not stolen. Our core idea is that a malicious accuser can deviate (without detection) from the specified MOR process by finding (transferable) adversarial examples that successfully serve as evidence against independent suspect models. To this end, we first generalize the procedures of common MOR schemes and show that, under this generalization, defending against false claims is as challenging as preventing (transferable) adversarial examples. Via systematic empirical evaluation we demonstrate that our false claim attacks always succeed in all prominent MOR schemes with realistic configurations, including against a real-world model: Amazon's Rekognition API.
翻译:深度神经网络(DNN)模型作为模型所有者的重要知识产权,构成其竞争优势。因此,开发保护技术以防范模型盗窃至关重要。模型所有权验证(Model Ownership Resolution, MOR)是一类能够威慑模型盗窃的技术。其方案允许指控方通过提交水印、指纹等证据,声称嫌疑模型系从其所拥有的源模型窃取或衍生而来。现有MOR方案大多侧重于抵御恶意嫌疑方,确保当嫌疑模型确为窃取模型时指控方能够胜诉。本文表明,文献中常见MOR方案存在另一同等重要但尚未充分探究的鲁棒性隐患:恶意指控方。我们展示了恶意指控方如何成功对未被窃取的独立嫌疑模型提出虚假声明。其核心思路是,恶意指控方可通过寻找(可迁移的)对抗样本作为证据(且不被察觉),偏离规定的MOR流程。为此,我们首先归纳了常见MOR方案的通用流程,并证明在此泛化框架下,防御虚假声明的困难程度与防范(可迁移)对抗样本相当。通过系统性实验评估,我们证明在所有主流MOR方案(含实际配置)中,包括针对真实世界模型(Amazon的Rekognition API)时,我们的虚假声明攻击始终成功。