Current large language models have dangerous capabilities, which are likely to become more problematic in the future. Activation steering techniques can be used to reduce risks from these capabilities. In this paper, we investigate the efficacy of activation steering for broad skills and multiple behaviours. First, by comparing the effects of reducing performance on general coding ability and Python-specific ability, we find that steering broader skills is competitive to steering narrower skills. Second, we steer models to become more or less myopic and wealth-seeking, among other behaviours. In our experiments, combining steering vectors for multiple different behaviours into one steering vector is largely unsuccessful. On the other hand, injecting individual steering vectors at different places in a model simultaneously is promising.
翻译:当前的大型语言模型具备危险能力,且这些能力在未来可能愈发成问题。激活引导技术可用于降低这些能力带来的风险。本文研究了激活引导在广泛技能与多种行为上的有效性。首先,通过对比降低通用编码能力与Python特定能力的效果,我们发现引导广泛技能与引导狭窄技能的效果相当。其次,我们引导模型在多种行为上(如减少或增加短视倾向与财富追求)发生变化。实验表明,将针对多种不同行为的引导向量合并为一个引导向量的方法效果不佳;另一方面,同时将多个引导向量分别注入模型的不同位置则展现出潜力。