In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on residual stream activations at each transformer block. We find that refusal is linearly decodable well before the final layer, indicating that safety-relevant behavior is represented in intermediate activations before output generation. To test whether this signal is actionable, we introduce Mechanistic AutoDAN, a probe-guided variant of AutoDAN that replaces full-model fitness evaluation with partial forward passes and probe-based scoring inside a genetic prompt search loop. Across the evaluated models, our method achieves attack success rates competitive with vanilla AutoDAN while reducing per-iteration search time by up to 72%, and probe-guided prompts match or exceed AutoDAN's cross-model transfer in several configurations. We further find that the usefulness of probe guidance increases with model scale. Our results show that refusal is not only observable at the output level, but is encoded as a structured and actionable signal in intermediate LLM activations.
翻译:本文通过在线性探针训练于每个Transformer块的残差流激活上,研究了是否能在解码前从大型语言模型(LLM)的中间激活中预测其拒绝行为。我们发现,远在最后一层之前,拒绝行为即可被线性解码,这表明与安全相关的行为在输出生成前的中间激活中已有表征。为测试该信号是否可行,我们提出Mechanistic AutoDAN——一种探针引导的AutoDAN变体,它在遗传提示搜索循环中用部分前向传播和基于探针的评分替代全模型适应性评估。在评估的模型上,我们的方法达到了与原始AutoDAN相当的攻击成功率,同时将每次迭代的搜索时间减少了高达72%,且在多种配置下,探针引导的提示在跨模型迁移方面匹配或超越了AutoDAN。我们进一步发现,探针引导的效用随模型规模增大而增强。我们的结果表明,拒绝不仅在输出层面可观测,更在LLM中间激活中被编码为一种结构化且可行的信号。