Recent work shows that in-context learning and optimization of in-context examples (ICE) can significantly improve the accuracy of large language models (LLMs) on a wide range of tasks, leading to an apparent consensus that ICE optimization is crucial for better performance. However, most of these studies assume a fixed or no instruction provided in the prompt. We challenge this consensus by investigating the necessity of optimizing ICE when task-specific instructions are provided and find that there are tasks for which it yields diminishing returns. In particular, using a diverse set of tasks and a systematically created instruction set with gradually added details, we find that as the prompt instruction becomes more detailed, the returns on ICE optimization diminish. To characterize this behavior, we introduce a task-specific metric called Normalized Invariability to Choice of Examples (NICE) that quantifies the learnability of tasks from a given instruction, and provides a heuristic that helps decide whether to optimize instructions or ICE for a new task. Given a task, the proposed metric can reliably predict the utility of optimizing ICE compared to using random ICE.
翻译:近期研究表明,上下文学习与上下文示例(ICE)优化能显著提升大语言模型(LLMs)在各类任务中的准确性,这使学界形成共识——ICE优化对提升性能至关重要。然而,多数研究假设提示中未包含指令或仅含固定指令。我们通过探究在提供任务专用指令时优化ICE的必要性,对这一共识提出质疑,发现部分任务中ICE优化的收益会递减。具体而言,通过采用多样化任务集和系统性构建的、逐步增加细节的指令集,我们发现当提示指令越详尽,ICE优化的收益衰减越明显。为量化这一特性,我们引入任务专用指标——标准化示例选择不变性(NICE),该指标可评估给定指令下任务的可学习性,并提供启发式方法辅助判断新任务中应优先优化指令还是ICE。针对特定任务,该指标能可靠预测相比随机选取ICE,优化ICE的实际效用。