Social alignment in AI systems aims to ensure that these models behave according to established societal values. However, unlike humans, who derive consensus on value judgments through social interaction, current language models (LMs) are trained to rigidly replicate their training corpus in isolation, leading to subpar generalization in unfamiliar scenarios and vulnerability to adversarial attacks. This work presents a novel training paradigm that permits LMs to learn from simulated social interactions. In comparison to existing methodologies, our approach is considerably more scalable and efficient, demonstrating superior performance in alignment benchmarks and human evaluations. This paradigm shift in the training of LMs brings us a step closer to developing AI systems that can robustly and accurately reflect societal norms and values.
翻译:人工智能系统中的社会一致性旨在确保这些模型遵循既定的社会价值观运行。然而,与人类通过社会互动达成价值判断共识不同,当前语言模型(LMs)在训练时被孤立地强制复制其训练语料库,导致在陌生场景中泛化能力不足且易受对抗性攻击。本研究提出一种新型训练范式,允许语言模型从模拟社会互动中学习。与现有方法相比,本方案在可扩展性和效率上具有显著优势,在一致性基准测试与人类评估中均展现出更优性能。这种语言模型训练范式的转变,使我们向开发能够稳健且精准反映社会规范与价值观的人工智能系统迈进了关键一步。