ChatGPT is a groundbreaking ``chatbot"--an AI interface built on a large language model that was trained on an enormous corpus of human text to emulate human conversation. Beyond its ability to converse in a plausible way, it has attracted attention for its ability to competently answer questions from the bar exam and from MBA coursework, and to provide useful assistance in writing computer code. These apparent abilities have prompted discussion of ChatGPT as both a threat to the integrity of higher education and conversely as a powerful teaching tool. In this work we present a preliminary analysis of how two versions of ChatGPT (ChatGPT3.5 and ChatGPT4) fare in the field of first-semester university physics, using a modified version of the Force Concept Inventory (FCI) to assess whether it can give correct responses to conceptual physics questions about kinematics and Newtonian dynamics. We demonstrate that, by some measures, ChatGPT3.5 can match or exceed the median performance of a university student who has completed one semester of college physics, though its performance is notably uneven and the results are nuanced. By these same measures, we find that ChatGPT4's performance is approaching the point of being indistinguishable from that of an expert physicist when it comes to introductory mechanics topics. After the completion of our work we became aware of Ref [1], which preceded us to publication and which completes an extensive analysis of the abilities of ChatGPT3.5 in a physics class, including a different modified version of the FCI. We view this work as confirming that portion of their results, and extending the analysis to ChatGPT4, which shows rapid and notable improvement in most, but not all respects.
翻译:ChatGPT是一种开创性的“聊天机器人”——一种基于大型语言模型的人工智能界面,该模型通过海量人类文本语料库训练而成,旨在模拟人类对话。除了能以合理的方式对话外,它还因能够胜任律师资格考试和MBA课程的问题回答,并在编写计算机代码时提供有效辅助而备受关注。这些显著能力引发了关于ChatGPT的讨论:它既是对高等教育诚信的威胁,也是一强大的教学工具。本文中,我们采用修改版力概念量表(FCI),对两种版本的ChatGPT(ChatGPT3.5和ChatGPT4)在大学第一学期物理领域的表现进行初步分析,评估其能否正确回答关于运动学和牛顿力学的概念性物理问题。我们证明,从某些指标看,ChatGPT3.5能够匹配甚至超越完成一学期大学物理课程的大学生中位数水平,但其表现显著不均衡且结果复杂。通过相同指标,我们发现ChatGPT4在初等力学主题上的表现已接近专家物理学家水平,几乎难以区分。在本工作完成后,我们得知了参考文献[1],该文献先于我们发表,并完成了对ChatGPT3.5在物理课堂中能力的全面分析(包括另一种修改版FCI)。我们认为本工作证实了他们结果中的部分内容,并将分析扩展到ChatGPT4,后者在大多数(尽管并非全部)方面显示出快速且显著的改进。