Adversarial attacks expose vulnerabilities of deep learning models by introducing minor perturbations to the input, which lead to substantial alterations in the output. Our research focuses on the impact of such adversarial attacks on sequence-to-sequence (seq2seq) models, specifically machine translation models. We introduce algorithms that incorporate basic text perturbation heuristics and more advanced strategies, such as the gradient-based attack, which utilizes a differentiable approximation of the inherently non-differentiable translation metric. Through our investigation, we provide evidence that machine translation models display robustness displayed robustness against best performed known adversarial attacks, as the degree of perturbation in the output is directly proportional to the perturbation in the input. However, among underdogs, our attacks outperform alternatives, providing the best relative performance. Another strong candidate is an attack based on mixing of individual characters.
翻译:对抗攻击通过在输入中引入微小扰动,导致输出发生显著变化,从而暴露深度学习模型的脆弱性。本研究聚焦于此类对抗攻击对序列到序列(seq2seq)模型(特别是机器翻译模型)的影响。我们提出了包含基本文本扰动启发式算法及更高级策略(如基于梯度的攻击)的算法,该策略利用翻译指标固有的不可微分性的可微近似。通过调查,我们提供了证据表明:机器翻译模型对已知最强对抗攻击展现出强健的鲁棒性,因为输出扰动程度与输入扰动程度成正比。然而,在弱攻击方法中,我们提出的攻击优于其他方案,实现了最佳相对性能。另一种强有力的候选方案是基于单个字符混合的扰动攻击。