Compositional generalization--understanding unseen combinations of seen primitives--is an essential reasoning capability in human intelligence. The AI community mainly studies this capability by fine-tuning neural networks on lots of training samples, while it is still unclear whether and how in-context learning--the prevailing few-shot paradigm based on large language models--exhibits compositional generalization. In this paper, we present CoFe, a test suite to investigate in-context compositional generalization. We find that the compositional generalization performance can be easily affected by the selection of in-context examples, thus raising the research question what the key factors are to make good in-context examples for compositional generalization. We study three potential factors: similarity, diversity and complexity. Our systematic experiments indicate that in-context examples should be structurally similar to the test case, diverse from each other, and individually simple. Furthermore, two strong limitations are observed: in-context compositional generalization on fictional words is much weaker than that on commonly used ones; it is still critical that the in-context examples should cover required linguistic structures, even though the backbone model has been pre-trained on large corpus. We hope our analysis would facilitate the understanding and utilization of in-context learning paradigm.
翻译:组合泛化——理解所见基元的未见组合——是人类智能中一项关键的推理能力。人工智能领域主要通过在海量训练样本上微调神经网络来研究此能力,但尚不清楚基于大语言模型的流行小样本范式——上下文学习——是否以及如何展现组合泛化。本文提出CoFe测试套件,用于探究上下文组合泛化。我们发现组合泛化性能易受上下文示例选择的影响,由此引发研究问题:为组合泛化构建良好上下文示例的关键因素是什么?我们考察了三个潜在因素:相似性、多样性和复杂性。系统性实验表明,上下文示例应与测试用例在结构上相似、彼此多样且各自简单。此外,观察到两个显著局限:虚构词汇的上下文组合泛化能力远弱于常见词汇;即便骨干模型已在大规模语料上预训练,上下文示例仍需覆盖所需的语言结构。我们希望本文分析能促进对上下文学习范式的理解与应用。