Large language models (LLMs) have become integral to our professional workflows and daily lives. Nevertheless, these machine companions of ours have a critical flaw: the huge amount of data which endows them with vast and diverse knowledge, also exposes them to the inevitable toxicity and bias. While most LLMs incorporate defense mechanisms to prevent the generation of harmful content, these safeguards can be easily bypassed with minimal prompt engineering. In this paper, we introduce the new Thoroughly Engineered Toxicity (TET) dataset, comprising manually crafted prompts designed to nullify the protective layers of such models. Through extensive evaluations, we demonstrate the pivotal role of TET in providing a rigorous benchmark for evaluation of toxicity awareness in several popular LLMs: it highlights the toxicity in the LLMs that might remain hidden when using normal prompts, thus revealing subtler issues in their behavior.
翻译:大规模语言模型(LLM)已成为我们专业工作流程和日常生活中的重要组成部分。然而,这些机器伴侣存在一个关键缺陷:海量的训练数据在赋予它们广泛多样的知识的同时,也使它们不可避免地接触到有害内容和偏见。尽管大多数LLM都内置了防御机制以防止生成有害内容,但这些防护措施往往可以通过简单的提示工程轻易绕过。本文提出了全新的系统性毒性工程(TET)数据集,该数据集包含人工精心设计的提示词,旨在消除这类模型的防护层。通过广泛评估,我们证明了TET在多个主流LLM的毒性意识评估中作为严格基准的关键作用:它能揭示使用常规提示词时可能隐藏的毒性问题,从而暴露模型行为中更细微的缺陷。