We explore the creative problem-solving capabilities of modern large language models (LLMs) in a constrained setting. The setting requires circumventing a cognitive bias known in psychology as ''functional fixedness'' to use familiar objects in innovative or unconventional ways. To this end, we create MacGyver, an automatically generated dataset consisting of 1,600 real-world problems that deliberately trigger functional fixedness and require thinking 'out-of-the-box'. We then present our collection of problems to both LLMs and humans to compare and contrast their problem-solving abilities. We show that MacGyver is challenging for both groups, but in unique and complementary ways. For example, humans typically excel in solving problems that they are familiar with but may struggle with tasks requiring domain-specific knowledge, leading to a higher variance. On the other hand, LLMs, being exposed to a variety of highly specialized knowledge, attempt broader problems but are prone to overconfidence and propose actions that are physically infeasible or inefficient. We also provide a detailed error analysis of LLMs, and demonstrate the potential of enhancing their problem-solving ability with novel prompting techniques such as iterative step-wise reflection and divergent-convergent thinking. This work provides insight into the creative problem-solving capabilities of humans and AI and illustrates how psychological paradigms can be extended into large-scale tasks for comparing humans and machines.
翻译:我们探索了现代大型语言模型在受限环境中的创造性问题解决能力。该环境要求规避心理学中称为“功能固着”的认知偏差,以创新或非传统方式使用熟悉物品。为此,我们构建了MacGyver,一个自动生成的包含1,600个真实世界问题的数据集,该数据集刻意诱发功能固着并需要“跳出框框”思考。随后我们将问题集合呈现给大型语言模型和人类,以比较和对比他们的问题解决能力。研究表明,MacGyver对两者均具有挑战性,但挑战方式独特且互补。例如,人类通常擅长解决熟悉问题,但在需要领域特定知识的任务上可能遇到困难,导致较高方差。相比之下,大型语言模型因接触多种高度专业化知识而尝试更广泛的问题,但易过度自信,并提出物理上不可行或低效的行动。我们还提供了大型语言模型的详细错误分析,并展示了通过创新提示技术(如迭代逐步反思和发散-收敛思维)提升其问题解决能力的潜力。这项工作揭示了人类与人工智能的创造性问题解决能力,并阐释了如何将心理学范式扩展到大规模任务中以比较人类与机器。