It has long been assumed that the sheer number of parameters in large language models (LLMs) drives in-context learning (ICL) capabilities, enabling remarkable performance improvements by leveraging task-specific demonstrations. Challenging this hypothesis, we introduce DEEP-ICL, a novel task Definition Enriched ExPert Ensembling methodology for ICL. DEEP-ICL explicitly extracts task definitions from given demonstrations and generates responses through learning task-specific examples. We argue that improvement from ICL does not directly rely on model size, but essentially stems from understanding task definitions and task-guided learning. Inspired by this, DEEP-ICL combines two 3B models with distinct roles (one for concluding task definitions and the other for learning task demonstrations) and achieves comparable performance to LLaMA2-13B. Furthermore, our framework outperforms conventional ICL by overcoming pretraining sequence length limitations, by supporting unlimited demonstrations. We contend that DEEP-ICL presents a novel alternative for achieving efficient few-shot learning, extending beyond the conventional ICL.
翻译:长期以来,大型语言模型(LLM)中庞大的参数数量被认为是驱动上下文学习(ICL)能力的关键,通过利用任务特定示例实现显著的性能提升。针对这一假设,我们提出了DEEP-ICL,一种新颖的基于任务定义增强的专家集成方法。DEEP-ICL通过给定示例显式提取任务定义,并基于学习任务特定样本来生成响应。我们认为,ICL的性能提升并非直接依赖于模型规模,而是本质上源于对任务定义的理解以及任务指导的学习过程。受此启发,DEEP-ICL结合了两个功能不同的3B模型(一个用于总结任务定义,另一个用于学习任务示例),并实现了与LLaMA2-13B相当的性能。此外,我们的框架通过支持无限数量的示例,克服了预训练序列长度限制,从而优于传统ICL方法。我们主张DEEP-ICL为实现高效少样本学习提供了一种超越传统ICL的新范式。