Recent advances have greatly increased the capabilities of large language models (LLMs), but our understanding of the models and their safety has not progressed as fast. In this paper we aim to understand LLMs deeper by studying their individual neurons. We build upon previous work showing large language models such as GPT-4 can be useful in explaining what each neuron in a language model does. Specifically, we analyze the effect of the prompt used to generate explanations and show that reformatting the explanation prompt in a more natural way can significantly improve neuron explanation quality and greatly reduce computational cost. We demonstrate the effects of our new prompts in three different ways, incorporating both automated and human evaluations.
翻译:近期进展极大地提升了大型语言模型(LLMs)的能力,但我们对模型本身及其安全性的理解并未同步加快。本文旨在通过研究单个神经元来更深入地理解LLMs。我们基于先前研究表明,诸如GPT-4等大型语言模型可用于解释语言模型中每个神经元的功能。具体而言,我们分析了用于生成解释的提示(prompt)的影响,并证明以更自然的方式重新格式化解释提示可以显著提升神经元解释质量,同时大幅降低计算成本。我们通过三种不同方式(涵盖自动评估与人工评估)验证了新提示的效果。