Projecting intermediate representations onto the vocabulary is an increasingly popular interpretation tool for transformer-based LLMs, also known as the logit lens. We propose a quantitative extension to this approach and define spectral filters on intermediate representations based on partitioning the singular vectors of the vocabulary embedding and unembedding matrices into bands. We find that the signals exchanged in the tail end of the spectrum are responsible for attention sinking (Xiao et al. 2023), of which we provide an explanation. We find that the loss of pretrained models can be kept low despite suppressing sizable parts of the embedding spectrum in a layer-dependent way, as long as attention sinking is preserved. Finally, we discover that the representation of tokens that draw attention from many tokens have large projections on the tail end of the spectrum.
翻译:将中间表示投影到词汇表上,已成为基于Transformer的大语言模型日益流行的解释工具,即所谓的"logit透镜"。我们对此方法提出定量扩展,通过将词汇嵌入矩阵和逆嵌入矩阵的奇异向量划分为频带,定义了中间表示上的频谱滤波器。研究发现,频谱尾端交换的信号是导致注意力沉降现象(Xiao et al. 2023)的原因,并对此提供了解释。我们发现在保持注意力沉降的前提下,即使以层级依赖方式抑制嵌入谱的显著部分,预训练模型的损失仍可维持在较低水平。最终发现,吸引众多令牌注意力的令牌表示在频谱尾端具有较大的投影分量。