The emergent capabilities of Large Language Models (LLMs) have made it crucial to align their values with those of humans. Current methodologies typically attempt alignment with a homogeneous human value and requires human verification, yet lack consensus on the desired aspect and depth of alignment and resulting human biases. In this paper, we propose A2EHV, an Automated Alignment Evaluation with a Heterogeneous Value system that (1) is automated to minimize individual human biases, and (2) allows assessments against various target values to foster heterogeneous agents. Our approach pivots on the concept of value rationality, which represents the ability for agents to execute behaviors that satisfy a target value the most. The quantification of value rationality is facilitated by the Social Value Orientation framework from social psychology, which partitions the value space into four categories to assess social preferences from agents' behaviors. We evaluate the value rationality of eight mainstream LLMs and observe that large models are more inclined to align neutral values compared to those with strong personal values. By examining the behavior of these LLMs, we contribute to a deeper understanding of value alignment within a heterogeneous value system.
翻译:大型语言模型(LLMs)的涌现能力使其与人类价值观的对齐变得至关重要。当前方法通常试图与同质化的人类价值观对齐,并依赖人工验证,但在对齐的目标方面、对齐深度以及由此产生的人类偏见方面缺乏共识。本文提出A2EHV——一种基于异构价值体系的自动化对齐评估方法,其特点是:(1)实现自动化以最小化个体人类偏见;(2)允许针对不同目标价值观进行评估,从而促进异构智能体的发展。我们的方法基于价值理性这一概念——它代表智能体执行最符合目标价值观行为的能力。价值理性的量化借助社会心理学中的社会价值取向框架,该框架将价值空间划分为四个类别,用于从智能体行为中评估社会偏好。我们对八个主流LLM的价值理性进行了评估,观察到大型模型相较于具有强烈个人价值观的模型更倾向于与中性价值观对齐。通过分析这些LLM的行为,我们为异构价值体系中的价值对齐提供了更深入的理解。