We propose a novel gradient-based attack against transformer-based language models that searches for an adversarial example in a continuous space of token probabilities. Our algorithm mitigates the gap between adversarial loss for continuous and discrete text representations by performing multi-step quantization in a quantization-compensation loop. Experiments show that our method significantly outperforms other approaches on various natural language processing (NLP) tasks.
翻译:我们提出了一种新颖的基于梯度的攻击方法,针对基于Transformer的语言模型,在连续的令牌概率空间中搜索对抗样本。我们的算法通过执行量化-补偿循环中的多步量化,弥合了连续与离散文本表示之间对抗损失的差距。实验表明,我们的方法在各种自然语言处理(NLP)任务上显著优于其他方法。