ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Safety is critical to the usage of large language models (LLMs). Multiple techniques such as data filtering and supervised fine-tuning have been developed to strengthen LLM safety. However, currently known techniques presume that corpora used for safety alignment of LLMs are solely interpreted by semantics. This assumption, however, does not hold in real-world applications, which leads to severe vulnerabilities in LLMs. For example, users of forums often use ASCII art, a form of text-based art, to convey image information. In this paper, we propose a novel ASCII art-based jailbreak attack and introduce a comprehensive benchmark Vision-in-Text Challenge (ViTC) to evaluate the capabilities of LLMs in recognizing prompts that cannot be solely interpreted by semantics. We show that five SOTA LLMs (GPT-3.5, GPT-4, Gemini, Claude, and Llama2) struggle to recognize prompts provided in the form of ASCII art. Based on this observation, we develop the jailbreak attack ArtPrompt, which leverages the poor performance of LLMs in recognizing ASCII art to bypass safety measures and elicit undesired behaviors from LLMs. ArtPrompt only requires black-box access to the victim LLMs, making it a practical attack. We evaluate ArtPrompt on five SOTA LLMs, and show that ArtPrompt can effectively and efficiently induce undesired behaviors from all five LLMs.

翻译：安全对于大型语言模型（LLMs）的使用至关重要。目前已开发出多种技术（如数据过滤和监督微调）来增强LLM的安全性。然而，现有技术均假设用于LLM安全对齐的语料库仅能通过语义进行解释。这一假设在实际应用中并不成立，从而导致LLM存在严重漏洞。例如，论坛用户常使用ASCII艺术（一种基于文本的艺术形式）来传达图像信息。本文提出了一种新颖的基于ASCII艺术的越狱攻击，并引入了全面的基准测试——视觉文本挑战（Vision-in-Text Challenge, ViTC），以评估LLM识别无法仅通过语义解释的提示的能力。研究表明，五种最先进的LLM（GPT-3.5、GPT-4、Gemini、Claude和Llama2）均难以识别以ASCII艺术形式呈现的提示。基于此发现，我们开发了越狱攻击ArtPrompt，利用LLM在识别ASCII艺术方面的性能缺陷来绕过安全措施，从而诱导LLM产生不良行为。ArtPrompt仅需对受害者LLM进行黑盒访问，具有实际攻击性。我们在五种最先进的LLM上评估了ArtPrompt，结果表明ArtPrompt能有效且高效地诱导所有五种LLM产生不良行为。

相关内容

大语言模型

关注 66

大语言模型是基于海量文本数据训练的深度学习模型。它不仅能够生成自然语言文本，还能够深入理解文本含义，处理各种自然语言任务，如文本摘要、问答、翻译等。2023年，大语言模型及其在人工智能领域的应用已成为全球科技研究的热点，其在规模上的增长尤为引人注目，参数量已从最初的十几亿跃升到如今的一万亿。参数量的提升使得模型能够更加精细地捕捉人类语言微妙之处，更加深入地理解人类语言的复杂性。在过去的一年里，大语言模型在吸纳新知识、分解复杂任务以及图文对齐等多方面都有显著提升。随着技术的不断成熟，它将不断拓展其应用范围，为人类提供更加智能化和个性化的服务，进一步改善人们的生活和生产方式。

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日