As a type of valuable intellectual property (IP), deep neural network (DNN) models have been protected by techniques like watermarking. However, such passive model protection cannot fully prevent model abuse. In this work, we propose an active model IP protection scheme, namely NNSplitter, which actively protects the model by splitting it into two parts: the obfuscated model that performs poorly due to weight obfuscation, and the model secrets consisting of the indexes and original values of the obfuscated weights, which can only be accessed by authorized users. NNSplitter uses the trusted execution environment to secure the secrets and a reinforcement learning-based controller to reduce the number of obfuscated weights while maximizing accuracy drop. Our experiments show that by only modifying 313 out of over 28 million (i.e., 0.001%) weights, the accuracy of the obfuscated VGG-11 model on Fashion-MNIST can drop to 10%. We also demonstrate that NNSplitter is stealthy and resilient against potential attack surfaces, including norm clipping and fine-tuning attacks.
翻译:作为一种有价值的知识产权,深度神经网络(DNN)模型已通过水印等技术受到保护。然而,这种被动式模型保护无法完全阻止模型滥用。本文提出了一种主动式模型知识产权保护方案——NNSplitter,该方案通过将模型分割为两部分实现对模型的主动保护:因权重混淆而导致性能低下的混淆模型,以及包含混淆权重索引和原始值的模型机密信息(仅授权用户可访问)。NNSplitter利用可信执行环境保护机密信息,并采用基于强化学习的控制器在最大化准确率下降的同时减少混淆权重数量。实验表明,仅需修改VGG-11模型超过2800万参数中的313个(即0.001%)权重,即可使该模型在Fashion-MNIST数据集上的准确率下降至10%。此外,我们证明了NNSplitter能够有效抵御范数裁剪和微调攻击等潜在攻击面,具有隐蔽性与鲁棒性。