Recent studies have revealed that NLP predictive models are vulnerable to adversarial attacks. Most existing studies focused on designing attacks to evaluate the robustness of NLP models in the English language alone. Literature has seen an increasing need for NLP solutions for other languages. We, therefore, ask one natural question: whether state-of-the-art (SOTA) attack methods generalize to other languages. This paper investigates how to adapt SOTA adversarial attack algorithms in English to the Chinese language. Our experiments show that attack methods previously applied to English NLP can generate high-quality adversarial examples in Chinese when combined with proper text segmentation and linguistic constraints. In addition, we demonstrate that the generated adversarial examples can achieve high fluency and semantic consistency by focusing on the Chinese language's morphology and phonology, which in turn can be used to improve the adversarial robustness of Chinese NLP models.
翻译:近期研究揭示,自然语言处理(NLP)预测模型易受对抗攻击影响。现有研究大多聚焦于设计针对英文NLP模型鲁棒性评估的攻击方法。随着文献显示对其他语言NLP解决方案的需求日益增长,我们因此提出一个自然问题:最先进(SOTA)攻击方法能否泛化至其他语言?本文探究如何将英文SOTA对抗攻击算法适配至中文。实验表明,结合恰当文本分割与语言约束后,先前应用于英文NLP的攻击方法可在中文中生成高质量对抗样本。此外,通过聚焦中文形态学与音韵学特征,我们证明所生成对抗样本能实现高流畅性与语义一致性,进而可用于提升中文NLP模型的对抗鲁棒性。