Large language models (LLMs) enable in-context learning (ICL) by conditioning on a few labeled training examples as a text-based prompt, eliminating the need for parameter updates and achieving competitive performance. In this paper, we demonstrate that factual knowledge is imperative for the performance of ICL in three core facets: the inherent knowledge learned in LLMs, the factual knowledge derived from the selected in-context examples, and the knowledge biases in LLMs for output generation. To unleash the power of LLMs in few-shot learning scenarios, we introduce a novel Knowledgeable In-Context Tuning (KICT) framework to further improve the performance of ICL: 1) injecting knowledge into LLMs during continual self-supervised pre-training, 2) judiciously selecting the examples for ICL with high knowledge relevance, and 3) calibrating the prediction results based on prior knowledge. We evaluate the proposed approaches on autoregressive models (e.g., GPT-style LLMs) over multiple text classification and question-answering tasks. Experimental results demonstrate that KICT substantially outperforms strong baselines and improves by more than 13% and 7% on text classification and question-answering tasks, respectively.
翻译:大型语言模型(LLMs)通过将少量标注训练样本作为文本提示进行条件化处理,无需参数更新即可实现上下文学习(ICL),并取得竞争性性能。本文证明,事实知识在ICL性能的三个核心层面不可或缺:LLMs中习得的内生知识、从所选上下文示例中提取的事实知识、以及LLMs在输出生成中的知识偏差。为激发LLMs在少样本学习场景中的潜力,我们提出创新框架"知识增强的上下文调优"(KICT)以进一步提升ICL性能:1)在持续自监督预训练阶段向LLMs注入知识,2)依据知识相关性审慎选择上下文学习示例,3)基于先验知识校准预测结果。我们在自回归模型(如GPT类LLMs)上,针对多项文本分类和问答任务评估所提方法。实验结果表明,KICT在文本分类和问答任务上分别显著超越强基线方法,提升幅度超过13%和7%。