Mechanistic interpretability aims to understand the behavior of neural networks by reverse-engineering their internal computations. However, current methods struggle to find clear interpretations of neural network activations because a decomposition of activations into computational features is missing. Individual neurons or model components do not cleanly correspond to distinct features or functions. We present a novel interpretability method that aims to overcome this limitation by transforming the activations of the network into a new basis - the Local Interaction Basis (LIB). LIB aims to identify computational features by removing irrelevant activations and interactions. Our method drops irrelevant activation directions and aligns the basis with the singular vectors of the Jacobian matrix between adjacent layers. It also scales features based on their importance for downstream computation, producing an interaction graph that shows all computationally-relevant features and interactions in a model. We evaluate the effectiveness of LIB on modular addition and CIFAR-10 models, finding that it identifies more computationally-relevant features that interact more sparsely, compared to principal component analysis. However, LIB does not yield substantial improvements in interpretability or interaction sparsity when applied to language models. We conclude that LIB is a promising theory-driven approach for analyzing neural networks, but in its current form is not applicable to large language models.
翻译:机制可解释性旨在通过逆向工程神经网络内部计算来理解其行为。然而,当前方法难以清晰解释神经网络激活,因为缺少将激活分解为计算特征的手段。单个神经元或模型组件并未与不同特征或功能形成明确对应关系。我们提出一种新颖的可解释性方法,通过将网络激活转换到新的基——局部交互基(LIB)来克服这一局限。LIB旨在通过移除无关激活和交互来识别计算特征。该方法丢弃无关激活方向,并将基与相邻层间雅可比矩阵的奇异向量对齐。它还根据特征对下游计算的重要性进行缩放,生成展示模型中所有计算相关特征及交互的交互图。我们在模加法和CIFAR-10模型上评估了LIB的有效性,发现相比主成分分析,它能识别更多计算相关特征且这些特征交互更稀疏。然而,当应用于语言模型时,LIB在可解释性或交互稀疏性方面未取得实质性改进。我们得出结论:LIB是一种有前景的理论驱动型神经网络分析方法,但当前形式尚不适用于大型语言模型。