Large language models based on transformers have achieved great empirical successes. However, as they are deployed more widely, there is a growing need to better understand their internal mechanisms in order to make them more reliable. These models appear to store vast amounts of knowledge from their training data, and to adapt quickly to new information provided in their context or prompt. We study how transformers balance these two types of knowledge by considering a synthetic setup where tokens are generated from either global or context-specific bigram distributions. By a careful empirical analysis of the training process on a simplified two-layer transformer, we illustrate the fast learning of global bigrams and the slower development of an "induction head" mechanism for the in-context bigrams. We highlight the role of weight matrices as associative memories, provide theoretical insights on how gradients enable their learning during training, and study the role of data-distributional properties.
翻译:基于Transformer的大语言模型已在实证上取得了巨大成功。然而,随着其广泛应用,为了提升模型的可靠性,迫切需要对内部机制进行更深入的理解。这些模型似乎能从训练数据中存储海量知识,并能快速适应上下文或提示中提供的新信息。我们通过考虑由全局双字母组分布或上下文特定双字母组分布生成的标记这一合成场景,研究了Transformer如何平衡这两类知识。通过对简化的两层Transformer训练过程进行细致的实证分析,我们揭示了全局双字母组的快速学习与上下文双字母组“归纳头”机制的缓慢形成过程。我们强调了权重矩阵作为联想记忆的作用,提供了关于梯度如何在训练期间促进其学习的理论见解,并研究了数据分布特性的影响。