Existing interpretability methods for Large Language Models (LLMs) predominantly capture linear directions or isolated features. This overlooks the high-dimensional, relational, and nonlinear geometry of model representations. We apply persistent homology (PH) to characterize how adversarial inputs reshape the geometry and topology of internal representation spaces of LLMs. This phenomenon, especially when considered across operationally different attack modes, remains poorly understood. We analyze six models (3.8B to 70B parameters) under two distinct attacks, indirect prompt injection and backdoor fine--tuning, and show that a consistent topological signature persists throughout. Adversarial inputs induce topological compression, where the latent space becomes structurally simpler, collapsing the latent space from varied, compact, small-scale features into fewer, dominant, large-scale ones. This signature is architecture-agnostic, emerges early in the network, and is highly discriminative across layers. By quantifying the shape of activation point clouds and neuron-level information flow, our framework reveals geometric invariants of representational change that complement existing linear interpretability methods.
翻译:现有大型语言模型(LLM)的可解释性方法主要捕捉线性方向或孤立特征,这忽视了模型表示的高维、关系性和非线性几何结构。我们应用持续同调(PH)来刻画对抗性输入如何重塑LLM内部表示空间的几何形态与拓扑结构。尤其当考虑操作上不同的攻击模式时,这一现象仍未被充分理解。我们分析了两种不同攻击(间接提示注入和后门微调)下的六个模型(参数量从3.8B到70B),并证明了一种一致的拓扑特征贯穿始终。对抗性输入会引发拓扑压缩,使得潜在空间在结构上趋于简单化——从多样化、紧凑的小尺度特征坍缩为少数主导性的大尺度特征。该特征与架构无关,在网络早期即出现,并具有高度的跨层可区分性。通过量化激活点云的形态和神经元级信息流,我们的框架揭示了表示变化的几何不变量,从而补充了现有的线性可解释性方法。