Historical emphasis on writing mastery has shifted with advances in generative AI, especially in scientific writing. This study analysed six AI chatbots for scholarly writing in humanities and archaeology. Using methods that assessed factual correctness and scientific contribution, ChatGPT-4 showed the highest quantitative accuracy, closely followed by ChatGPT-3.5, Bing, and Bard. However, Claude 2 and Aria scored considerably lower. Qualitatively, all AIs exhibited proficiency in merging existing knowledge, but none produced original scientific content. Inter-estingly, our findings suggest ChatGPT-4 might represent a plateau in large language model size. This research emphasizes the unique, intricate nature of human research, suggesting that AI's emulation of human originality in scientific writing is challenging. As of 2023, while AI has transformed content generation, it struggles with original contributions in humanities. This may change as AI chatbots continue to evolve into LLM-powered software.
翻译:历史上对写作能力的重视随着生成式人工智能的进步而发生转变,尤其是在科学写作领域。本研究分析了六款AI聊天机器人在人文学科和考古学学术写作中的表现。采用评估事实准确性和科学贡献的方法,ChatGPT-4在量化准确性上表现最高,ChatGPT-3.5、Bing和Bard紧随其后。然而,Claude 2和Aria的得分明显较低。在定性方面,所有AI均展现出整合现有知识的能力,但均未能产生原创性科学内容。有趣的是,我们的研究结果表明,ChatGPT-4可能代表了大型语言模型规模的一个平台期。这项研究强调了人类研究独特且复杂的本质,表明AI在科学写作中模仿人类原创性具有挑战性。截至2023年,尽管AI已经改变了内容生成方式,但在人文学科的原创性贡献方面仍存在困难。随着AI聊天机器人持续演变为基于大型语言模型的软件,这一情况未来可能发生变化。