Large Language Models (LLMs) have the impressive ability to perform in-context learning (ICL) from only a few examples, but the success of ICL varies widely from task to task. Thus, it is important to quickly determine whether ICL is applicable to a new task, but directly evaluating ICL accuracy can be expensive in situations where test data is expensive to annotate -- the exact situations where ICL is most appealing. In this paper, we propose the task of ICL accuracy estimation, in which we predict the accuracy of an LLM when doing in-context learning on a new task given only unlabeled test data for that task. To perform ICL accuracy estimation, we propose a method that trains a meta-model using LLM confidence scores as features. We compare our method to several strong accuracy estimation baselines on a new benchmark that covers 4 LLMs and 3 task collections. The meta-model improves over all baselines across 8 out of 12 settings and achieves the same estimation performance as directly evaluating on 40 collected labeled test examples per task. At the same time, no existing approach provides an accurate and reliable ICL accuracy estimation in every setting, highlighting the need for better ways to measure the uncertainty of LLM predictions.
翻译:大语言模型(LLMs)具有仅通过少量示例进行上下文学习(ICL)的卓越能力,但ICL的成功率因任务而异。因此,快速判断ICL是否适用于新任务至关重要,然而在测试数据标注成本高昂的场景中——这正是ICL最具吸引力的应用场景——直接评估ICL准确率可能代价高昂。本文提出ICL准确率估算任务,即仅利用新任务的未标注测试数据,预测大语言模型在执行上下文学习时的准确率。为实现ICL准确率估算,我们提出一种元模型方法,该方法以大语言模型的置信度分数作为特征进行训练。我们在涵盖4种大语言模型和3个任务集合的新基准上,将所提方法与多个强基线模型进行对比。结果表明,该元模型在12个设置中的8个上优于所有基线,且其估算性能等同于对每个任务直接评估40个已标注测试样本。与此同时,现有方法尚未能在所有场景下提供准确可靠的ICL准确率估算,这凸显了改进大语言模型预测不确定性测量方法的必要性。