To what extent can language alone give rise to complex concepts, or is embodied experience essential? Recent advancements in large language models (LLMs) offer fresh perspectives on this question. Although LLMs are trained on restricted modalities, they exhibit human-like performance in diverse psychological tasks. Our study compared representations of 4,442 lexical concepts between humans and ChatGPTs (GPT-3.5 and GPT-4) across multiple dimensions, including five key domains: emotion, salience, mental visualization, sensory, and motor experience. We identify two main findings: 1) Both models strongly align with human representations in non-sensorimotor domains but lag in sensory and motor areas, with GPT-4 outperforming GPT-3.5; 2) GPT-4's gains are associated with its additional visual learning, which also appears to benefit related dimensions like haptics and imageability. These results highlight the limitations of language in isolation, and that the integration of diverse modalities of inputs leads to a more human-like conceptual representation.
翻译:语言本身能在多大程度上产生复杂的概念?具身经验是否不可或缺?大型语言模型的最新进展为这一问题提供了全新视角。尽管LLM仅基于受限模态进行训练,却在多种心理学任务中展现出类人表现。本研究从情感、显著度、心理可视化、感觉和运动体验五个关键领域出发,比较了人类与ChatGPT(GPT-3.5与GPT-4)对4,442个词汇概念的表征。我们得出两个主要发现:1)两个模型在非感觉运动领域与人类表征高度一致,但在感觉和运动领域存在差距,且GPT-4表现优于GPT-3.5;2)GPT-4的进步与其额外的视觉学习相关,这种学习似乎也促进了触觉感知和可意象性等相关维度的改善。这些结果揭示了纯语言的局限性,并表明整合多模态输入能够产生更接近人类的表征形式。