Mechanistic interpretability aims to understand model behaviors in terms of specific, interpretable features, often hypothesized to manifest as low-dimensional subspaces of activations. Specifically, recent studies have explored subspace interventions (such as activation patching) as a way to simultaneously manipulate model behavior and attribute the features behind it to given subspaces. In this work, we demonstrate that these two aims diverge, potentially leading to an illusory sense of interpretability. Counterintuitively, even if a subspace intervention makes the model's output behave as if the value of a feature was changed, this effect may be achieved by activating a dormant parallel pathway leveraging another subspace that is causally disconnected from model outputs. We demonstrate this phenomenon in a distilled mathematical example, in two real-world domains (the indirect object identification task and factual recall), and present evidence for its prevalence in practice. In the context of factual recall, we further show a link to rank-1 fact editing, providing a mechanistic explanation for previous work observing an inconsistency between fact editing performance and fact localization. However, this does not imply that activation patching of subspaces is intrinsically unfit for interpretability. To contextualize our findings, we also show what a success case looks like in a task (indirect object identification) where prior manual circuit analysis informs an understanding of the location of a feature. We explore the additional evidence needed to argue that a patched subspace is faithful.
翻译:机械可解释性旨在通过特定、可解释的特征来理解模型行为,这些特征常被假设表现为激活向量的低维子空间。具体而言,近期研究探索了子空间干预(如激活修补)作为同时操控模型行为并将行为背后的特征归因到给定子空间的方法。本研究表明,这两个目标可能相互背离,从而产生一种可解释性的错觉。反直觉的是,即使子空间干预使模型输出表现如同某个特征的值被改变,这种效果也可能是通过激活一条利用另一与模型输出因果解耦的子空间的休眠并行通路实现的。我们在一个简化数学示例、两个实际领域(间接宾语识别任务和事实回忆)中展示了这一现象,并提供了其在实践中普遍存在的证据。在事实回忆的背景下,我们还展示了其与秩1事实编辑之间的联系,为先前观察到的编辑性能与事实定位不一致的研究提供了机制解释。然而,这并不意味着子空间激活修补本身不适合可解释性。为了定位我们的发现,我们还展示了一个成功案例——在间接宾语识别任务中,先前的手动电路分析已揭示特征位置的理解。我们探讨了证明修补子空间可靠性所需的额外证据。