As the use of large language models (LLMs) increases within society, as does the risk of their misuse. Appropriate safeguards must be in place to ensure LLM outputs uphold the ethical standards of society, highlighting the positive role that artificial intelligence technologies can have. Recent events indicate ethical concerns around conventionally trained LLMs, leading to overall unsafe user experiences. This motivates our research question: how do we ensure LLM alignment? In this work, we introduce a test suite of unique prompts to foster the development of aligned LLMs that are fair, safe, and robust. We show that prompting LLMs at every step of the development pipeline, including data curation, pre-training, and fine-tuning, will result in an overall more responsible model. Our test suite evaluates outputs from four state-of-the-art language models: GPT-3.5, GPT-4, OPT, and LLaMA-2. The assessment presented in this paper highlights a gap between societal alignment and the capabilities of current LLMs. Additionally, implementing a test suite such as ours lowers the environmental overhead of making models safe and fair.
翻译:随着大型语言模型(LLMs)在社会中的应用日益广泛,其被滥用的风险也随之增加。必须建立适当的保障措施以确保LLM输出符合社会的伦理标准,从而彰显人工智能技术所能发挥的积极作用。近期事件表明,传统训练的LLM存在伦理隐忧,导致整体用户体验不安全。这引出了我们的研究问题:如何确保LLM的对齐性?本文中,我们引入了一套独特的提示测试集,以促进开发公平、安全且稳健的对齐LLM。我们证明,在开发流程的每个阶段(包括数据整理、预训练和微调)进行提示测试,将产生更负责任的模型。我们的测试集评估了四种最先进语言模型的输出:GPT-3.5、GPT-4、OPT 和 LLaMA-2。本文的评估揭示了社会对齐要求与当前LLM能力之间的差距。此外,实施此类测试集能降低实现模型安全与公平所需的环境开销。