The potential of large language models (LLMs) to reason like humans has been a highly contested topic in Machine Learning communities. However, the reasoning abilities of humans are multifaceted and can be seen in various forms, including analogical, spatial and moral reasoning, among others. This fact raises the question whether LLMs can perform equally well across all these different domains. This research work aims to investigate the performance of LLMs on different reasoning tasks by conducting experiments that directly use or draw inspirations from existing datasets on analogical and spatial reasoning. Additionally, to evaluate the ability of LLMs to reason like human, their performance is evaluted on more open-ended, natural language questions. My findings indicate that LLMs excel at analogical and moral reasoning, yet struggle to perform as proficiently on spatial reasoning tasks. I believe these experiments are crucial for informing the future development of LLMs, particularly in contexts that require diverse reasoning proficiencies. By shedding light on the reasoning abilities of LLMs, this study aims to push forward our understanding of how they can better emulate the cognitive abilities of humans.
翻译:大型语言模型(LLM)像人类一样推理的潜力一直是机器学习社区中备受争议的话题。然而,人类的推理能力是多方面的,可以以多种形式体现,包括类比推理、空间推理和道德推理等。这一事实引发了疑问:LLM是否能在所有这些不同领域中同样出色地表现?本研究旨在通过直接使用或借鉴现有类比推理和空间推理数据集的实验,探究LLM在不同推理任务上的表现。此外,为了评估LLM像人类一样推理的能力,本研究还在更开放的自然语言问题上对其性能进行了评估。研究发现表明,LLM在类比推理和道德推理方面表现出色,但在空间推理任务上难以达到同样熟练的水平。我认为这些实验对于指导LLM的未来发展至关重要,特别是在需要多样化推理能力的场景中。通过揭示LLM的推理能力,本研究旨在推动我们对其如何更好地模拟人类认知能力的理解。