In this paper, we introduce a Bayesian learning model to understand the behavior of Large Language Models (LLMs). We explore the optimization metric of LLMs, which is based on predicting the next token, and develop a novel model grounded in this principle. Our approach involves constructing an ideal generative text model represented by a multinomial transition probability matrix with a prior, and we examine how LLMs approximate this matrix. We discuss the continuity of the mapping between embeddings and multinomial distributions, and present the Dirichlet approximation theorem to approximate any prior. Additionally, we demonstrate how text generation by LLMs aligns with Bayesian learning principles and delve into the implications for in-context learning, specifically explaining why in-context learning emerges in larger models where prompts are considered as samples to be updated. Our findings indicate that the behavior of LLMs is consistent with Bayesian Learning, offering new insights into their functioning and potential applications.
翻译:本文提出了一种贝叶斯学习模型,用于理解大语言模型(LLMs)的行为机制。我们深入探究了LLMs基于下一个词元预测的优化指标,并以此为基础开发了一种新颖的理论模型。该方法通过构建具有先验分布的多项式转移概率矩阵作为理想生成文本模型,进而考察LLMs对该矩阵的逼近过程。我们论证了嵌入向量与多项式分布之间映射的连续性,并提出了狄利克雷逼近定理以实现任意先验分布的近似。此外,我们揭示了LLMs的文本生成过程与贝叶斯学习原理的一致性,并深入探讨了上下文学习的本质机制——具体阐释了为何在更大规模模型中,将提示视为待更新样本时能够涌现出上下文学习能力。研究结果表明,LLMs的行为模式与贝叶斯学习框架高度契合,为其功能机制与潜在应用提供了全新的理论视角。