How do language models "think"? This paper formulates a probabilistic cognitive model called the bounded pragmatic speaker, which can characterize the operation of different variations of language models. Specifically, we demonstrate that large language models fine-tuned with reinforcement learning from human feedback (Ouyang et al., 2022) embody a model of thought that conceptually resembles a fast-and-slow model (Kahneman, 2011), which psychologists have attributed to humans. We discuss the limitations of reinforcement learning from human feedback as a fast-and-slow model of thought and propose avenues for expanding this framework. In essence, our research highlights the value of adopting a cognitive probabilistic modeling approach to gain insights into the comprehension, evaluation, and advancement of language models.
翻译:语言模型如何“思考”?本文提出了一个称为有界语用说话者的概率认知模型,该模型能够表征不同变体的语言模型的运作方式。具体而言,我们证明,通过人类反馈的强化学习(Ouyang等人,2022)微调的大型语言模型体现了一种思维模型,该模型在概念上类似于心理学家归因于人类的“快-慢”模型(Kahneman,2011)。我们讨论了强化学习从人类反馈中作为快-慢思维模型的局限性,并提出了扩展该框架的途径。本质上,我们的研究强调了采用认知概率建模方法对于深入理解、评估和推进语言模型的价值。