What has ChatGPT read? The origins of archaeological citations used by a generative artificial intelligence application

The public release of ChatGPT has resulted in considerable publicity and has led to wide-spread discussion of the usefulness and capabilities of generative AI language models. Its ability to extract and summarise data from textual sources and present them as human-like contextual responses makes it an eminently suitable tool to answer questions users might ask. This paper tested what archaeological literature appears to have been included in ChatGPT's training phase. While ChatGPT offered seemingly pertinent references, a large percentage proved to be fictitious. Using cloze analysis to make inferences on the sources 'memorised' by a generative AI model, this paper was unable to prove that ChatGPT had access to the full texts of the genuine references. It can be shown that all references provided by ChatGPT that were found to be genuine have also been cited on Wikipedia pages. This strongly indicates that the source base for at least some of the data is found in those pages. The implications of this in relation to data quality are discussed.

翻译：ChatGPT的公开上线引发了广泛关注，并促使学界对生成式AI语言模型的实用性与能力展开深入讨论。该模型能够从文本资料中提取并总结数据，并以类人化的情境响应呈现，使其成为解答用户问题的理想工具。本文测试了哪些考古学文献可能被纳入ChatGPT的训练阶段。尽管ChatGPT提供了看似相关的参考文献，但其中相当大比例被证实为虚构。通过使用完形分析推断生成式AI模型"记忆"的来源，本研究无法证明ChatGPT能够获取真实参考文献的全文。研究表明，所有经查证为真实的ChatGPT所引用的参考文献，均曾出现在维基百科页面中。这强烈表明，至少部分数据源来自这些维基百科页面。本文进一步讨论了这一发现对数据质量的影响。

相关内容

ChatGPT

关注 258

ChatGPT（全名：Chat Generative Pre-trained Transformer），美国OpenAI 研发的聊天机器人程序 [1] ，于2022年11月30日发布。ChatGPT是人工智能技术驱动的自然语言处理工具，它能够通过学习和理解人类的语言来进行对话，还能根据聊天的上下文进行互动，真正像人类一样来聊天交流，甚至能完成撰写邮件、视频脚本、文案、翻译、代码，写论文任务。 [1] https://openai.com/blog/chatgpt/

《生成式模型: 变分自编码器与扩散模型》，75页ppt，Google DeepMind科学家Ruiqi Gao

专知会员服务

66+阅读 · 2023年6月10日

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日