Machine learning security has recently become a prominent topic in the natural language processing (NLP) area. The existing black-box adversarial attack suffers prohibitively from the high model querying complexity, resulting in easily being captured by anti-attack monitors. Meanwhile, how to eliminate redundant model queries is rarely explored. In this paper, we propose a query-efficient approach BufferSearch to effectively attack general intelligent NLP systems with the minimal number of querying requests. In general, BufferSearch makes use of historical information and conducts statistical test to avoid incurring model queries frequently. Numerically, we demonstrate the effectiveness of BufferSearch on various benchmark text-classification experiments by achieving the competitive attacking performance but with a significant reduction of query quantity. Furthermore, BufferSearch performs multiple times better than competitors within restricted query budget. Our work establishes a strong benchmark for the future study of query-efficiency in NLP adversarial attacks.
翻译:机器学习安全近来已成为自然语言处理领域的一个突出话题。现有的黑盒对抗攻击因模型查询复杂度极高而代价高昂,导致容易被反攻击监测器捕获。与此同时,如何消除冗余模型查询这一问题鲜有探索。本文提出一种查询高效的方法BufferSearch,能够以最少的查询请求次数有效攻击通用智能NLP系统。总体而言,BufferSearch利用历史信息并进行统计检验,以避免频繁引发模型查询。通过数值实验,我们在多个基准文本分类任务上展示了BufferSearch的有效性,其攻击性能具有竞争力,但查询数量显著减少。此外,在受限的查询预算内,BufferSearch的表现比竞品好数倍。我们的工作为NLP对抗攻击中查询效率的未来研究建立了强有力的基准。