Recent studies on adversarial examples expose vulnerabilities of natural language processing (NLP) models. Existing techniques for generating adversarial examples are typically driven by deterministic heuristic rules that are agnostic to the optimal adversarial examples, a strategy that often results in attack failures. To this end, this research proposes Fraud's Bargain Attack (FBA) which utilizes a novel randomization mechanism to enlarge the search space and enables high-quality adversarial examples to be generated with high probabilities. FBA applies the Metropolis-Hasting sampler, a member of Markov Chain Monte Carlo samplers, to enhance the selection of adversarial examples from all candidates proposed by a customized stochastic process that we call the Word Manipulation Process (WMP). WMP perturbs one word at a time via insertion, removal or substitution in a contextual-aware manner. Extensive experiments demonstrate that FBA outperforms the state-of-the-art methods in terms of both attack success rate and imperceptibility.
翻译:近期关于对抗性样本的研究揭示了自然语言处理(NLP)模型的脆弱性。现有生成对抗性样本的技术通常由确定性启发式规则驱动,这些规则无法感知最优对抗性样本,导致攻击往往以失败告终。为此,本研究提出欺诈竞价攻击(FBA),该攻击利用一种新颖的随机化机制扩大搜索空间,从而以高概率生成高质量对抗性样本。FBA采用马尔可夫链蒙特卡洛采样器中的梅特罗波利斯-黑斯廷斯采样器,从我们自定义的随机过程——词语操控过程(WMP)所提出的所有候选方案中增强对抗性样本的选取。WMP通过上下文感知的方式,每次对单个词语进行插入、删除或替换扰动。广泛实验表明,FBA在攻击成功率与不可感知性方面均优于现有最先进方法。