Subliminal learning describes a student language model inheriting a behavioral bias by fine-tuning on seemingly innocuous data generated by a biased teacher model. Prior work has begun to characterize this phenomenon but leaves open questions about the scope of signals it can transfer, the mechanisms that explain it, and the precision with which a bias can be encoded by seemingly unrelated data. We tackle all three problems by introducing subliminal steering, a variant of subliminal learning in which the teacher's bias is implemented not via a system prompt, as in prior work, but through a steering vector trained to maximize the likelihood of a set of target samples. First, we show that subliminal steering transfers complex multi-word biases, whereas prior work focused on single-word preferences, demonstrating a large scope of subliminally transferrable signals. Second, we provide mechanistic evidence that subliminal learning transfers not only the target behavioral bias, but also the steering vector itself, localized to the layers at which the teacher was steered. Finally, we show that the bias is encoded with surprising precision. We train a new steering vector directly on the subliminally-laden dataset and find that it attains high cosine similarity with the original vector.
翻译:潜意识学习描述了学生语言模型通过微调由有偏教师模型生成的看似无害数据,继承行为偏差的现象。已有研究开始刻画这一现象,但对其可传递信号的范围、解释机制以及偏差如何通过看似无关的数据被精确编码等问题仍未解决。我们通过引入"潜意识引导"——一种教师偏差不通过系统提示(如先前研究)实现,而是通过优化使目标样本集似然最大化的引导向量来实施的潜意识学习变体——对这三个问题展开研究。首先,我们发现潜意识引导能传递复杂的多词偏差,而先前研究仅关注单词偏好,这揭示了潜意识可传递信号的广泛范围。其次,我们提供了机制性证据,表明潜意识学习不仅传递目标行为偏差,还传递引导向量本身,且该向量定位于教师被引导的特定层。最后,我们证明了这种偏差以惊人的精度被编码。通过在带有潜意识标记的数据集上直接训练新引导向量,我们发现该向量与原向量具有高余弦相似度。