While large language models based on the transformer architecture have demonstrated remarkable in-context learning (ICL) capabilities, understandings of such capabilities are still in an early stage, where existing theory and mechanistic understanding focus mostly on simple scenarios such as learning simple function classes. This paper takes initial steps on understanding ICL in more complex scenarios, by studying learning with representations. Concretely, we construct synthetic in-context learning problems with a compositional structure, where the label depends on the input through a possibly complex but fixed representation function, composed with a linear function that differs in each instance. By construction, the optimal ICL algorithm first transforms the inputs by the representation function, and then performs linear ICL on top of the transformed dataset. We show theoretically the existence of transformers that approximately implement such algorithms with mild depth and size. Empirically, we find trained transformers consistently achieve near-optimal ICL performance in this setting, and exhibit the desired dissection where lower layers transforms the dataset and upper layers perform linear ICL. Through extensive probing and a new pasting experiment, we further reveal several mechanisms within the trained transformers, such as concrete copying behaviors on both the inputs and the representations, linear ICL capability of the upper layers alone, and a post-ICL representation selection mechanism in a harder mixture setting. These observed mechanisms align well with our theory and may shed light on how transformers perform ICL in more realistic scenarios.
翻译:尽管基于Transformer架构的大语言模型展现出卓越的上下文学习能力,但对此能力的理解仍处于早期阶段,现有理论和机制性研究主要聚焦于简单函数类学习等场景。本文通过研究表征学习,初步探索了更复杂场景下的上下文学习机制。具体而言,我们构造了具有组合结构的合成上下文学习问题:其中标签依赖于通过复杂但固定的表征函数映射后的输入,该表征函数与每个实例中不同的线性函数相结合。通过设计,最优的上下文学习算法首先通过表征函数变换输入,然后在变换后的数据集上执行线性上下文学习。我们从理论上证明了存在适当深度和规模的Transformer能近似实现此类算法。实验表明,经过训练的Transformer在此设置中持续达到接近最优的上下文学习性能,并展现出预期的分层结构:底层变换数据集,顶层执行线性上下文学习。通过广泛的探针实验和新型拼接实验,我们进一步揭示了训练后Transformer中的多种机制,包括输入和表征的具体复制行为、顶层独立的线性上下文学习能力,以及更困难的混合设置中的上下文学习后表征选择机制。这些观测机制与我们的理论高度吻合,或可为理解Transformer在更现实场景中的上下文学习机制提供启示。