Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and system prompts appear to shape much of their behaviour. But models can also adopt different personas which have radically different preferences. How is this implemented internally? Does each persona run on its own preference machinery, or is something shared underneath? We train linear probes on residual-stream activations of Gemma-3-27B and Qwen-3.5-122B to predict revealed pairwise task choices, and identify a genuine preference vector: it tracks the model's preferences as they shift across a range of prompts and situations, and on Gemma-3-27B steering along it causally controls pairwise choice. This preference representation is largely shared across personas: a probe trained on the helpful assistant predicts and steers the choices of qualitatively different personas, including an evil persona whose preferences anti-correlate with those of the Assistant.
翻译:大型语言模型(LLMs)可被视为具有偏好:它们可靠地选择特定任务与输出而非其他,且由后训练与系统提示塑造的偏好似乎主导了其大部分行为。但模型也能采用具有根本不同偏好的多种角色。这种偏好如何在内部实现?每个角色是否运行于自身偏好机制,还是存在某种共享基础?我们对Gemma-3-27B与Qwen-3.5-122B的残差流激活进行线性探针训练,以预测所揭示的成对任务选择,并识别出真实偏好向量:该向量能追踪模型偏好随提示与情境范围的变化,且在Gemma-3-27B中沿此向量进行引导可因果控制成对选择。这种偏好表征在角色间基本共享:在有益助手上训练的探针可预测并引导包含邪恶角色(其偏好与助手反相关)等性质不同角色的选择行为。