Whilst Large Language Models (LLMs) have a weak to moderate ability to score published journal articles for research quality, they have not been compared with individual expert reviewers. It is also unknown whether quality scores from ChatGPT based on PDFs can improve on those from titles and abstracts through a deeper evaluation. To address both issues, this article uses expert scores (98 internal departmental ratings for UK Unit of Assessment [UoA] 3 Allied Health Professions, 44 for UoA13 Architecture, and 58 library and information science articles from UoA34), comparing them against ChatGPT-5.4 scores from both title/abstract and PDF inputs. For UoA3, individual reviewer scores were also compared against each other and ChatGPT-5.4. The rank correlations with expert scores are almost all statistically significantly positive, but differences between the correlations are mostly not, despite weakly suggesting that ChatGPT-5.4 can be more reliable than individual reviewers for UoA3. Moreover, whilst ChatGPT-5.4 provides more detailed evaluations of PDFs than of titles/abstracts, its score predictions do not seem to improve. Thus, whilst the results broadly confirm the value of ChatGPT scores for ranking academic documents, its apparently deeper evaluative comments on PDFs are misleading in the sense of not translating to improved score predictions.


翻译:暂无翻译

0
下载
关闭预览

相关内容

天大最新《大型语言模型评估》全面综述,111页pdf
专知会员服务
89+阅读 · 2023年10月31日
《利用 ChatGPT 实现高效事实核查》
专知会员服务
48+阅读 · 2023年10月25日
近期语音类前沿论文
深度学习每日摘要
14+阅读 · 2019年3月17日
disentangled-representation-papers
CreateAMind
26+阅读 · 2018年9月12日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
何为协作武器?
专知会员服务
6+阅读 · 9月1日
《理解认知战:超越信息》
专知会员服务
10+阅读 · 9月1日
美国战争部在GenAI.mil上推出OpenAI的ChatGPT Mil
专知会员服务
8+阅读 · 8月31日
人工智能赋能军事维护:重新定义国防战备
专知会员服务
4+阅读 · 8月31日
《美陆军野战手册(2026年):特种部队》
专知会员服务
6+阅读 · 8月31日
受限仓库多智能体取送中的动态安全等待点选择
相关VIP内容
天大最新《大型语言模型评估》全面综述,111页pdf
专知会员服务
89+阅读 · 2023年10月31日
《利用 ChatGPT 实现高效事实核查》
专知会员服务
48+阅读 · 2023年10月25日
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员