Although large language models (LLMs) exhibit remarkable capacity to leverage in-context demonstrations, it is still unclear to what extent they can learn new concepts or facts from ground-truth labels. To address this question, we examine the capacity of instruction-tuned LLMs to follow in-context concept guidelines for sentence labeling tasks. We design guidelines that present different types of factual and counterfactual concept definitions, which are used as prompts for zero-shot sentence classification tasks. Our results show that although concept definitions consistently help in task performance, only the larger models (with 70B parameters or more) have limited ability to work under counterfactual contexts. Importantly, only proprietary models such as GPT-3.5 and GPT-4 can recognize nonsensical guidelines, which we hypothesize is due to more sophisticated alignment methods. Finally, we find that Falcon-180B-chat is outperformed by Llama-2-70B-chat is most cases, which indicates that careful fine-tuning is more effective than increasing model scale. Altogether, our simple evaluation method reveals significant gaps in concept understanding between the most capable open-source language models and the leading proprietary APIs.
翻译:尽管大型语言模型(LLMs)展现出利用上下文示例的卓越能力,但其从真实标注中学习新概念或事实的程度仍不明确。为探究此问题,我们考察了指令微调LLMs在句子标注任务中遵循上下文概念准则的能力。我们设计了呈现不同类型事实性与反事实性概念定义的准则,并将其作为零样本句子分类任务的提示。研究结果表明:虽然概念定义能持续提升任务表现,但仅有较大规模模型(70B参数及以上)在反事实语境中具备有限的处理能力。值得注意的是,只有GPT-3.5和GPT-4等专有模型能够识别无意义的准则,我们推测这源于其更复杂的对齐方法。最后,我们发现Falcon-180B-chat在多数情况下表现不及Llama-2-70B-chat,这表明精细的微调比单纯扩大模型规模更为有效。总体而言,我们提出的简易评估方法揭示了当前最先进的开源语言模型与领先专有API在概念理解层面存在显著差距。