The growing popularity of Deep Neural Networks, which often require computationally expensive training and access to a vast amount of data, calls for accurate authorship verification methods to deter unlawful dissemination of the models and identify the source of the leak. In DNN watermarking the owner may have access to the full network (white-box) or only be able to extract information from its output to queries (black-box), but a watermarked model may include both approaches in order to gather sufficient evidence to then gain access to the network. Although there has been limited research in white-box watermarking that considers traitor tracing, this problem is yet to be explored in the black-box scenario. In this paper, we propose a black-and-white-box watermarking method for DNN classifiers that opens the door to collusion-resistant traitor tracing in black-box, exploiting the properties of Tardos codes, and making it possible to identify the source of the leak before access to the model is granted. While experimental results show that the method can successfully identify traitors, even when further attacks have been performed, we also discuss its limitations and open problems for traitor tracing in black-box.
翻译:深度神经网络的普及(其训练通常需要高昂计算成本及海量数据访问权限)催生了精准的作者验证方法,用以遏制模型的非法传播并追溯泄露源头。在深度神经网络水印中,模型所有者可能拥有完整网络访问权限(白盒场景),或仅能通过输出查询提取信息(黑盒场景)。为收集足够证据进而获取网络访问权限,被水印模型可能同时融合这两种方案。尽管已有少量白盒水印研究涉及叛逆者追踪问题,但黑盒场景下的相关探索仍属空白。本文提出一种面向深度神经网络分类器的黑盒-白盒水印方法:通过利用Tardos码特性,在模型访问权限授予前即可从黑盒侧实现抗合谋的叛逆者追踪。实验表明,即便遭遇后续攻击,该方法仍能成功识别叛逆者;同时本文探讨了该方法的局限性及黑盒场景下叛逆者追踪面临的开放性问题。