Language model alignment has become an important component of AI safety, allowing safe interactions between humans and language models, by enhancing desired behaviors and inhibiting undesired ones. It is often done by tuning the model or inserting preset aligning prompts. Recently, representation engineering, a method which alters the model's behavior via changing its representations post-training, was shown to be effective in aligning LLMs (Zou et al., 2023a). Representation engineering yields gains in alignment oriented tasks such as resistance to adversarial attacks and reduction of social biases, but was also shown to cause a decrease in the ability of the model to perform basic tasks. In this paper we study the tradeoff between the increase in alignment and decrease in helpfulness of the model. We propose a theoretical framework which provides bounds for these two quantities, and demonstrate their relevance empirically. Interestingly, we find that while the helpfulness generally decreases, it does so quadratically with the norm of the representation engineering vector, while the alignment increases linearly with it, indicating a regime in which it is efficient to use representation engineering. We validate our findings empirically, and chart the boundaries to the usefulness of representation engineering for alignment.
翻译:语言模型对齐已成为AI安全的重要组成部分,通过增强期望行为并抑制非期望行为,实现人类与语言模型之间的安全交互。该过程通常通过调整模型或插入预设对齐提示实现。近期研究表明,表征工程——一种通过改变模型训练后表征来调整行为的方法——能有效对齐大语言模型(Zou等,2023a)。表征工程虽能在对抗攻击抵抗、社会偏见减少等对齐导向任务中取得增益,但也会导致模型执行基础任务的能力下降。本文研究模型对齐度提升与有用性下降之间的权衡关系。我们提出一个理论框架,为这两个指标设定边界,并通过实验验证其相关性。有趣的是,我们发现尽管有用性整体呈下降趋势,但其下降幅度与表征工程向量的范数呈二次关系,而对齐度则与该范数呈线性增长,这表明存在表征工程高效应用的区间。我们通过实证验证研究发现,并描绘了表征工程用于对齐的有效边界。