Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effective than optimization-free ones. However, current approaches to fine-tuned SVs suffer from two limitations. First, they require careful selection of steering factors on a per-SV basis to balance steering effectiveness and generation quality at inference time. Second, they operate as full-sequence SVs (FSSVs), which can sacrifice generation quality regardless of factor selection due to excessive intervention on the model generation process. To address the first limitation, we propose joint training of steering factors and directions, such that post-hoc factor selection is no longer required. Using neural network scaling theory, we find that moderately large initialization sizes and learning rates for steering factors are essential for stability and efficiency of joint training. To tackle the second limitation, we draw inspiration from representation fine-tuning and introduce Prompt-only SV (PrOSV), an SV that intervenes only on a few prompt tokens. Our empirical results show that PrOSV outperforms traditional FSSVs on AxBench when using our joint training scheme. We also find that PrOSV achieves a better tradeoff between general model utility and adversarial robustness than FSSV.
翻译:近年来,操控向量(SVs)已成为一种有效且轻量级的方法,用于调控大型语言模型(LLMs)的行为。其中,经过微调的操控向量比无需优化的方法更高效。然而,当前微调操控向量的方法存在两大局限:首先,在推理阶段需为每个操控向量仔细选择操控因子,以平衡操控效果与生成质量;其次,它们作为全序列操控向量(FSSVs)运行,无论因子选择如何,都可能因过度干预模型生成过程而牺牲生成质量。为解决第一个局限,我们提出联合训练操控因子与方向的方法,从而无需后续因子选择。基于神经网络缩放理论,我们发现适中的操控因子初始规模与学习率对联合训练的稳定性与效率至关重要。针对第二个局限,我们从表示微调中汲取灵感,引入仅提示干预的操控向量(PrOSV),该向量仅对少量提示词元进行干预。实验结果表明,结合我们的联合训练方案,PrOSV在AxBench基准上优于传统FSSVs。此外,我们发现与FSSV相比,PrOSV能在通用模型效用与对抗鲁棒性之间实现更优权衡。