How do language models "think"? This paper formulates a probabilistic cognitive model called the bounded pragmatic speaker, which can characterize the operation of different variations of language models. Specifically, we demonstrate that large language models fine-tuned with reinforcement learning from human feedback (Ouyang et al., 2022) embody a model of thought that conceptually resembles a fast-and-slow model (Kahneman, 2011), which psychologists have attributed to humans. We discuss the limitations of reinforcement learning from human feedback as a fast-and-slow model of thought and propose avenues for expanding this framework. In essence, our research highlights the value of adopting a cognitive probabilistic modeling approach to gain insights into the comprehension, evaluation, and advancement of language models.
翻译:语言模型如何“思考”?本文提出了一种名为有界实用说话者的概率认知模型,该模型能够表征不同变体语言模型的运作机制。具体而言,我们证明了通过人类反馈强化学习微调的大型语言模型(Ouyang等人,2022)体现了一种概念上类似于心理学家归因于人类的快慢思考模型(Kahneman,2011)的思维模式。我们讨论了人类反馈强化学习作为快慢思考模型的局限性,并提出了扩展该框架的途径。本质上,我们的研究强调了采用认知概率建模方法对理解、评估和推进语言模型的重要价值。