Measuring the creativity of large language models (LLMs) is essential for designing methods that can improve creativity and for enhancing our scientific understanding of this ability. To accomplish this, it has become common in recent years to administer tests of human creativity to LLMs. Although these tests provide a convenient and fully automated way to score "creativity," their validity as measures of machine creativity has not been established, and these tests already have limited validity as predictors of human creativity. To address this problem, we conduct the first large-scale, systematic study assessing the effectiveness of human creativity tests for predicting the creative achievement of LLMs across three target constructs: creative writing, divergent thinking, and scientific ideation. We find that the Divergent Association Task (DAT) and the Conditional DAT are the best predictors of creative writing and divergent thinking, respectively, but that test effectiveness varies significantly by construct, and no single test predicts all constructs well. Moreover, contrary to popular belief, no existing test reliably predicts scientific ideation ability. Motivated by this problem, we introduce the Divergent Remote Association Test (DRAT), a vocabulary-space test that assesses both convergent and divergent thinking in a single instrument. The DRAT is the first and only creativity test for LLMs that is a significant predictor of scientific ideation ability, demonstrating robustness across major design choices. Furthermore, the performance gain of the DRAT is not recoverable from any linear combination of the Divergent Association Task and the Remote Associates Test, indicating that assessing divergent and convergent thinking in the same test is essential to reliably predicting scientific ideation ability.


翻译:评估大型语言模型(LLMs)的创造力,对于设计提升其创造力的方法以及增强我们对该能力的科学理解至关重要。为此,近年来,将人类创造力测试应用于LLMs已成为一种普遍做法。尽管这些测试为评价"创造力"提供了便捷且全自动的评分方式,但作为机器创造力测量工具的有效性尚未得到验证,且这些测试在预测人类创造力方面本就存在局限性。针对这一问题,我们开展了首次大规模系统性研究,评估人类创造力测试在预测LLMs三项目标构念(创意写作、发散性思维与科学构想能力)的创造性成果方面的有效性。研究发现,发散联想任务(DAT)和条件性发散联想任务(Conditional DAT)分别是预测创意写作与发散性思维的最佳指标,但测试有效性在不同构念间存在显著差异,且没有任何单一测试能全面预测所有构念。此外,与普遍认知相反,现有测试均无法可靠预测科学构想能力。基于这一困境,我们提出了发散性远距离联想测试(DRAT)——一种在单一工具中同时评估聚合性思维与发散性思维的词汇空间测试。DRAT是首个且唯一能显著预测科学构想能力的LLM创造力测试,并在主要设计选择中展现出稳健性。更重要的是,任何线性组合(发散联想任务与远距联想测试的组合)均无法复现DRAT的性能增益,这表明在同一测试中同时评估发散性与聚合性思维,对可靠预测科学构想能力至关重要。

0
下载
关闭预览

相关内容

评估大语言模型在科学发现中的作用
专知会员服务
19+阅读 · 2025年12月19日
【斯坦福博士论文】大语言模型的AI辅助评估
专知会员服务
31+阅读 · 2025年3月30日
174页!《大语言模型》最新综述:能力与局限性分析
专知会员服务
65+阅读 · 2025年1月12日
大型语言模型的知识蒸馏综述:方法、评估与应用
专知会员服务
81+阅读 · 2024年7月4日
「大型语言模型评测」综述
专知会员服务
70+阅读 · 2024年3月30日
天大最新《大型语言模型评估》全面综述,111页pdf
专知会员服务
89+阅读 · 2023年10月31日
「知识增强预训练语言模型」最新研究综述
专知
18+阅读 · 2022年11月18日
自然语言处理中的语言模型预训练方法
PaperWeekly
14+阅读 · 2018年10月21日
NLP通用模型诞生?一个模型搞定十大自然语言常见任务
人工智能头条
10+阅读 · 2018年6月29日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
28+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Arxiv
26+阅读 · 2024年2月9日
Arxiv
17+阅读 · 2023年9月26日
Arxiv
21+阅读 · 2023年7月12日
Arxiv
25+阅读 · 2023年6月23日
Arxiv
12+阅读 · 2023年5月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
9+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
5+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
4+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
评估大语言模型在科学发现中的作用
专知会员服务
19+阅读 · 2025年12月19日
【斯坦福博士论文】大语言模型的AI辅助评估
专知会员服务
31+阅读 · 2025年3月30日
174页!《大语言模型》最新综述:能力与局限性分析
专知会员服务
65+阅读 · 2025年1月12日
大型语言模型的知识蒸馏综述:方法、评估与应用
专知会员服务
81+阅读 · 2024年7月4日
「大型语言模型评测」综述
专知会员服务
70+阅读 · 2024年3月30日
天大最新《大型语言模型评估》全面综述,111页pdf
专知会员服务
89+阅读 · 2023年10月31日
相关基金
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
28+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员