Search engines are vulnerable to attacks against indexing and searching via text encoding manipulation. By imperceptibly perturbing text using uncommon encoded representations, adversaries can control results across search engines for specific search queries. We demonstrate that this attack is successful against two major commercial search engines - Google and Bing - and one open source search engine - Elasticsearch. We further demonstrate that this attack is successful against LLM chat search including Bing's GPT-4 chatbot and Google's Bard chatbot. We also present a variant of the attack targeting text summarization and plagiarism detection models, two ML tasks closely tied to search. We provide a set of defenses against these techniques and warn that adversaries can leverage these attacks to launch disinformation campaigns against unsuspecting users, motivating the need for search engine maintainers to patch deployed systems.
翻译:搜索引擎在面对通过文本编码操纵进行的索引和搜索攻击时存在脆弱性。通过使用不常见的编码表示对文本进行难以察觉的扰动,攻击者可以控制特定搜索查询在不同搜索引擎上的结果。我们证明,该攻击对两大商业搜索引擎——谷歌(Google)和必应(Bing)——以及一个开源搜索引擎——Elasticsearch——均有效。我们还进一步证明,该攻击对大型语言模型聊天搜索(包括必应的GPT-4聊天机器人和谷歌的Bard聊天机器人)同样有效。此外,我们提出了一种针对文本摘要和抄袭检测模型(两个与搜索密切相关的机器学习任务)的攻击变体。我们提供了一套针对这些技术的防御措施,并警告攻击者可利用这些攻击对毫无戒备的用户发起虚假信息宣传活动,这凸显了搜索引擎维护者修补已部署系统的必要性。