In-context learning (ICL) has become an effective solution for few-shot learning in natural language processing. Past work has found that, during this process, representations of the last prompt token are utilized to store task reasoning procedures, thereby explaining the working mechanism of in-context learning. In this paper, we seek to locate and analyze other task-encoding tokens whose representations store task reasoning procedures. Supported by experiments that ablate the representations of different token types, we find that template and stopword tokens are the most prone to be task-encoding tokens. In addition, we demonstrate experimentally that lexical cues, repetition, and text formats are the main distinguishing characteristics of these tokens. Our work provides additional insights into how large language models (LLMs) leverage task reasoning procedures in ICL and suggests that future work may involve using task-encoding tokens to improve the computational efficiency of LLMs at inference time and their ability to handle long sequences.
翻译:上下文学习已成为自然语言处理中少样本学习的有效解决方案。先前研究发现,在此过程中,最后一个提示标记的表示被用于存储任务推理过程,从而解释了上下文学习的工作机制。本文旨在定位并分析其他任务编码标记,这些标记的表示同样存储了任务推理过程。通过消融不同类型标记表示的实验支持,我们发现模板标记和停用词标记最易成为任务编码标记。此外,我们通过实验证明,词汇线索、重复和文本格式是这些标记的主要区分特征。本研究为理解大型语言模型如何在上下文学习中利用任务推理过程提供了新的见解,并表明未来工作可利用任务编码标记在推理时提升大型语言模型的计算效率及其处理长序列的能力。