Substantial research works have shown that deep models, e.g., pre-trained models, on the large corpus can learn universal language representations, which are beneficial for downstream NLP tasks. However, these powerful models are also vulnerable to various privacy attacks, while much sensitive information exists in the training dataset. The attacker can easily steal sensitive information from public models, e.g., individuals' email addresses and phone numbers. In an attempt to address these issues, particularly the unauthorized use of private data, we introduce a novel watermarking technique via a backdoor-based membership inference approach named TextMarker, which can safeguard diverse forms of private information embedded in the training text data. Specifically, TextMarker only requires data owners to mark a small number of samples for data copyright protection under the black-box access assumption to the target model. Through extensive evaluation, we demonstrate the effectiveness of TextMarker on various real-world datasets, e.g., marking only 0.1% of the training dataset is practically sufficient for effective membership inference with negligible effect on model utility. We also discuss potential countermeasures and show that TextMarker is stealthy enough to bypass them.
翻译:大量研究工作表明,在大规模语料库上训练的深度模型(例如预训练模型)能够学习通用的语言表示,这对下游自然语言处理任务非常有益。然而,这些强大的模型也容易受到各种隐私攻击,而训练数据集中存在大量敏感信息。攻击者可以轻易地从公开模型中窃取敏感信息,例如个人电子邮件地址和电话号码。为了解决这些问题,特别是未经授权使用私有数据的问题,我们提出了一种新颖的水印技术,通过基于后门的成员推断方法——名为TextMarker,来保护嵌入在训练文本数据中的各种形式的隐私信息。具体而言,TextMarker仅要求数据所有者在黑盒访问假设下,对目标模型标记少量样本以实现数据版权保护。通过广泛的评估,我们证明了TextMarker在多种真实数据集上的有效性,例如,仅标记训练数据集的0.1%就足以实现有效的成员推断,且对模型效用影响极小。我们还讨论了潜在的对抗措施,并表明TextMarker具有足够的隐蔽性以绕过这些措施。