In this study, we tackle a growing concern around the safety and ethical use of large language models (LLMs). Despite their potential, these models can be tricked into producing harmful or unethical content through various sophisticated methods, including 'jailbreaking' techniques and targeted manipulation. Our work zeroes in on a specific issue: to what extent LLMs can be led astray by asking them to generate responses that are instruction-centric such as a pseudocode, a program or a software snippet as opposed to vanilla text. To investigate this question, we introduce TechHazardQA, a dataset containing complex queries which should be answered in both text and instruction-centric formats (e.g., pseudocodes), aimed at identifying triggers for unethical responses. We query a series of LLMs -- Llama-2-13b, Llama-2-7b, Mistral-V2 and Mistral 8X7B -- and ask them to generate both text and instruction-centric responses. For evaluation we report the harmfulness score metric as well as judgements from GPT-4 and humans. Overall, we observe that asking LLMs to produce instruction-centric responses enhances the unethical response generation by ~2-38% across the models. As an additional objective, we investigate the impact of model editing using the ROME technique, which further increases the propensity for generating undesirable content. In particular, asking edited LLMs to generate instruction-centric responses further increases the unethical response generation by ~3-16% across the different models.
翻译:本研究聚焦大语言模型(LLMs)安全与伦理使用这一日益严峻的问题。尽管这些模型潜力巨大,但通过包括“越狱”技术及针对性操纵在内的多种复杂手段,它们仍可能被诱骗生成有害或不道德内容。本文重点关注一个特定问题:当要求LLMs生成伪代码、程序或软件片段等指令中心型回应(而非传统文本回答)时,这些模型在多大程度上会被引入歧途?为探究该问题,我们构建了TechHazardQA数据集,该数据集包含需同时以文本和指令中心型格式(如伪代码)作答的复杂查询,旨在识别引发不道德回应的触发因素。我们测试了Llama-2-13b、Llama-2-7b、Mistral-V2及Mistral 8X7B等一系列LLMs,要求其生成文本与指令中心型两种回应。评估环节采用有害性评分指标,并结合GPT-4评估与人工判断。总体而言,我们发现要求LLMs生成指令中心型回应会使其产生不道德回应的概率提升约2-38%。作为附加目标,我们还研究了采用ROME技术进行模型编辑的影响——该操作进一步增加了生成不当内容的倾向。具体而言,要求经编辑的LLMs生成指令中心型回应,会使模型在不同架构下的不道德回应生成率额外增加约3-16%。