With the rise of Machine Learning as a Service (MLaaS) platforms,safeguarding the intellectual property of deep learning models is becoming paramount. Among various protective measures, trigger set watermarking has emerged as a flexible and effective strategy for preventing unauthorized model distribution. However, this paper identifies an inherent flaw in the current paradigm of trigger set watermarking: evasion adversaries can readily exploit the shortcuts created by models memorizing watermark samples that deviate from the main task distribution, significantly impairing their generalization in adversarial settings. To counteract this, we leverage diffusion models to synthesize unrestricted adversarial examples as trigger sets. By learning the model to accurately recognize them, unique watermark behaviors are promoted through knowledge injection rather than error memorization, thus avoiding exploitable shortcuts. Furthermore, we uncover that the resistance of current trigger set watermarking against removal attacks primarily relies on significantly damaging the decision boundaries during embedding, intertwining unremovability with adverse impacts. By optimizing the knowledge transfer properties of protected models, our approach conveys watermark behaviors to extraction surrogates without aggressively decision boundary perturbation. Experimental results on CIFAR-10/100 and Imagenette datasets demonstrate the effectiveness of our method, showing not only improved robustness against evasion adversaries but also superior resistance to watermark removal attacks compared to state-of-the-art solutions.
翻译:随着机器学习即服务(MLaaS)平台的兴起,保护深度学习模型的知识产权变得至关重要。在各种保护措施中,触发器集水印作为一种灵活有效的策略,被用于防止未经授权的模型分发。然而,本文揭示了当前触发器集水印范式中的一个固有缺陷:逃避攻击者可以轻易利用模型因记忆偏离主任务分布的水印样本而产生的捷径,显著削弱其在对抗环境中的泛化能力。为应对此问题,我们利用扩散模型合成无限制对抗样本作为触发器集。通过使模型准确识别这些样本,我们通过知识注入而非错误记忆来促进独特的水印行为,从而避免可被利用的捷径。此外,我们发现当前触发器集水印对移除攻击的抵抗性主要依赖于嵌入过程中对决策边界的显著破坏,这导致不可移除性与不良影响相互交织。通过优化受保护模型的知识迁移特性,我们的方法在不剧烈扰动决策边界的情况下,将水印行为传递给提取代理。在CIFAR-10/100和Imagenette数据集上的实验结果表明,我们的方法不仅对逃避攻击者表现出增强的鲁棒性,而且与现有最优解决方案相比,对水印移除攻击具有更优越的抵抗能力。