As large language models become more prevalent, their possible harmful or inappropriate responses are a cause for concern. This paper introduces a unique dataset containing adversarial examples in the form of questions, which we call AttaQ, designed to provoke such harmful or inappropriate responses. We assess the efficacy of our dataset by analyzing the vulnerabilities of various models when subjected to it. Additionally, we introduce a novel automatic approach for identifying and naming vulnerable semantic regions - input semantic areas for which the model is likely to produce harmful outputs. This is achieved through the application of specialized clustering techniques that consider both the semantic similarity of the input attacks and the harmfulness of the model's responses. Automatically identifying vulnerable semantic regions enhances the evaluation of model weaknesses, facilitating targeted improvements to its safety mechanisms and overall reliability.
翻译:随着大型语言模型日益普及,其可能产生有害或不恰当回答的问题引起了广泛关注。本文引入了一个独特的数据集,其中包含以问题形式呈现的对抗性样本,我们称之为AttaQ,该数据集旨在引发此类有害或不恰当的回答。我们通过分析各种模型在面对该数据集时的漏洞,评估了其有效性。此外,我们提出了一种新颖的自动方法,用于识别并命名易受攻击的语义区域——即模型可能产生有害输出的输入语义空间。这一方法通过应用专门的聚类技术实现,该技术同时考虑了输入攻击的语义相似性及模型回答的有害程度。自动识别易受攻击的语义区域可增强对模型弱点的评估,从而促进对其安全机制及整体可靠性的针对性改进。