As a core cognitive skill that enables the transferability of information across domains, analogical reasoning has been extensively studied for both humans and computational models. However, while cognitive theories of analogy often focus on narratives and study the distinction between surface, relational, and system similarities, existing work in natural language processing has a narrower focus as far as relational analogies between word pairs. This gap brings a natural question: can state-of-the-art large language models (LLMs) detect system analogies between narratives? To gain insight into this question and extend word-based relational analogies to relational system analogies, we devise a comprehensive computational framework that operationalizes dominant theories of analogy, using narrative elements to create surface and system mappings. Leveraging the interplay between these mappings, we create a binary task and benchmark for Analogical Reasoning on Narratives (ARN), covering four categories of far (cross-domain)/near (within-domain) analogies and disanalogies. We show that while all LLMs can largely recognize near analogies, even the largest ones struggle with far analogies in a zero-shot setting, with GPT4.0 scoring below random. Guiding the models through solved examples and chain-of-thought reasoning enhances their analogical reasoning ability. Yet, since even in the few-shot setting, the best model only performs halfway between random and humans, ARN opens exciting directions for computational analogical reasoners.
翻译:作为实现跨领域信息可迁移性的核心认知技能,类比推理已在人类与计算模型中受到广泛研究。然而,尽管类比的认知理论常聚焦于叙事,并区分表面相似性、关系相似性与系统相似性,自然语言处理领域现有研究却更局限于词对间的关系类比。这一鸿沟引出一个自然问题:当前最先进的大语言模型能否检测叙事间的系统类比?为深入探究此问题并将基于词汇的关系类延伸至关系系统类比,我们设计了一个全面的计算框架,该框架将主流的类比理论操作化,利用叙事元素创建表面映射与系统映射。通过利用这些映射间的相互作用,我们构建了一个二分类任务及基准——叙事类比推理(ARN),涵盖远(跨领域)/近(领域内)类比与反类比四大类别。研究表明,虽然所有大语言模型基本能识别近类比,但即便是最大的模型在零样本设定下仍难以应对远类比(GPT4.0得分低于随机水平)。通过引导模型使用示例解答与思维链推理,其类比推理能力得到提升。然而,即使在少样本设定下,最优模型的表现也仅介于随机水平与人类之间。由此,ARN为计算类比推理器的研究开辟了激动人心的方向。