Embedding as a Service (EaaS) has become a widely adopted solution, which offers feature extraction capabilities for addressing various downstream tasks in Natural Language Processing (NLP). Prior studies have shown that EaaS can be prone to model extraction attacks; nevertheless, this concern could be mitigated by adding backdoor watermarks to the text embeddings and subsequently verifying the attack models post-publication. Through the analysis of the recent watermarking strategy for EaaS, EmbMarker, we design a novel CSE (Clustering, Selection, Elimination) attack that removes the backdoor watermark while maintaining the high utility of embeddings, indicating that the previous watermarking approach can be breached. In response to this new threat, we propose a new protocol to make the removal of watermarks more challenging by incorporating multiple possible watermark directions. Our defense approach, WARDEN, notably increases the stealthiness of watermarks and empirically has been shown effective against CSE attack.
翻译:嵌入即服务(Embedding as a Service, EaaS)已成为广泛采用的解决方案,其为自然语言处理(Natural Language Processing, NLP)中的各类下游任务提供特征提取能力。先前研究表明,EaaS易受模型提取攻击;然而,通过向文本嵌入添加后门水印并在发布后验证攻击模型,这一隐患可得到缓解。通过分析近期EaaS水印策略EmbMarker,我们设计了一种新型CSE(聚类、选择、消除)攻击,该攻击在保持嵌入高实用性的同时移除了后门水印,表明先前的水印方法可被攻破。为应对这一新威胁,我们提出了一种新协议,通过整合多个可能的水印方向来增加水印移除的难度。我们的防御方法WARDEN显著提升了水印的隐蔽性,并且经验证对CSE攻击具有有效性。