Large Language Models (LLMs) are now commonplace in conversation applications. However, their risks of misuse for generating harmful responses have raised serious societal concerns and spurred recent research on LLM conversation safety. Therefore, in this survey, we provide a comprehensive overview of recent studies, covering three critical aspects of LLM conversation safety: attacks, defenses, and evaluations. Our goal is to provide a structured summary that enhances understanding of LLM conversation safety and encourages further investigation into this important subject. For easy reference, we have categorized all the studies mentioned in this survey according to our taxonomy, available at: https://github.com/niconi19/LLM-conversation-safety.
翻译:大型语言模型(LLM)在对话应用中已变得普遍。然而,它们被滥用于生成有害回复的风险引发了严重的社会关切,并推动了近期关于LLM对话安全的研究。因此,在本综述中,我们系统地梳理了近年来的研究,涵盖LLM对话安全的三个关键方面:攻击、防御与评估。我们的目标是提供一份结构化总结,以加深对LLM对话安全的理解,并鼓励对该重要课题的进一步探索。为便于查阅,我们已根据分类体系将所有提及的研究进行归类,详见:https://github.com/niconi19/LLM-conversation-safety。