DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space Smoothing

Large web-scale datasets have driven the rapid advancement of pre-trained language models (PLMs), but unauthorized data usage has raised serious copyright concerns. Existing dataset ownership verification (DOV) methods typically assume that watermarks remain stable during inference; however, this assumption often fails under natural noise and adversary-crafted perturbations. We propose the first certified dataset ownership verification method for PLMs under a gray-box setting (i.e., the defender can only query the suspicious model but is aware of its input representation module), based on dual-space smoothing (i.e., DSSmoothing). To address the challenges of text discreteness and semantic sensitivity, DSSmoothing introduces continuous perturbations in the embedding space to capture semantic robustness and applies controlled token reordering in the permutation space to capture sequential robustness. DSSmoothing consists of two stages: in the first stage, triggers are collaboratively embedded in both spaces to generate norm-constrained and robust watermarked datasets; in the second stage, randomized smoothing is applied in both spaces during verification to compute the watermark robustness (WR) of suspicious models and statistically compare it with the principal probability (PP) values of a set of benign models. Theoretically, DSSmoothing provides provable robustness guarantees for dataset ownership verification by ensuring that WR consistently exceeds PP under bounded dual-space perturbations. Extensive experiments on multiple representative web datasets demonstrate that DSSmoothing achieves stable and reliable verification performance and exhibits robustness against potential adaptive attacks. Our code is available at https://github.com/NcepuQiaoTing/DSSmoothing.

翻译：大规模网络数据集推动了预训练语言模型（PLMs）的快速发展，但未经授权的数据使用引发了严重的版权问题。现有的数据集所有权验证（DOV）方法通常假设水印在推理过程中保持稳定；然而，这一假设在自然噪声和对抗性扰动下往往失效。我们提出了首个在灰盒设置下（即防御者仅能查询可疑模型但知晓其输入表示模块）适用于PLMs的认证数据集所有权验证方法，该方法基于双空间平滑（即DSSmoothing）。为应对文本离散性和语义敏感性的挑战，DSSmoothing在嵌入空间引入连续扰动以捕获语义鲁棒性，并在置换空间应用受控的词元重排序以捕获序列鲁棒性。DSSmoothing包含两个阶段：第一阶段，在双空间协同嵌入触发器以生成范数约束且鲁棒的水印数据集；第二阶段，在验证过程中对双空间应用随机平滑，以计算可疑模型的水印鲁棒性（WR），并将其与一组良性模型的主概率（PP）值进行统计比较。理论上，DSSmoothing通过确保在有界双空间扰动下WR始终超过PP，为数据集所有权验证提供了可证明的鲁棒性保证。在多个代表性网络数据集上的大量实验表明，DSSmoothing实现了稳定可靠的验证性能，并对潜在的自适应攻击表现出鲁棒性。我们的代码发布于 https://github.com/NcepuQiaoTing/DSSmoothing。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

158页《大型语言模型数据集》全面综述，444个数据集涵盖预训练、指令微调、偏好、评估等，附中英文版

专知会员服务

155+阅读 · 2024年3月1日

[ACL2023]领域适配器混合：将领域知识解耦并注入到预训练语言模型的记忆中

专知会员服务

33+阅读 · 2023年6月11日

【ICML2023】调整语言模型作为增强少样本学习的训练数据生成器

专知会员服务

32+阅读 · 2023年5月19日