Massive open online courses (MOOCs) unlock the doors to free education for anyone around the globe with access to a computer and the internet. Despite this democratization of learning, the massive enrollment in these courses means it is almost impossible for one instructor to assess every student's writing assignment. As a result, peer grading, often guided by a straightforward rubric, is the method of choice. While convenient, peer grading often falls short in terms of reliability and validity. In this study, using 18 distinct settings, we explore the feasibility of leveraging large language models (LLMs) to replace peer grading in MOOCs. Specifically, we focus on two state-of-the-art LLMs: GPT-4 and GPT-3.5, across three distinct courses: Introductory Astronomy, Astrobiology, and the History and Philosophy of Astronomy. To instruct LLMs, we use three different prompts based on a variant of the zero-shot chain-of-thought (Zero-shot-CoT) prompting technique: Zero-shot-CoT combined with instructor-provided correct answers; Zero-shot-CoT in conjunction with both instructor-formulated answers and rubrics; and Zero-shot-CoT with instructor-offered correct answers and LLM-generated rubrics. Our results show that Zero-shot-CoT, when integrated with instructor-provided answers and rubrics, produces grades that are more aligned with those assigned by instructors compared to peer grading. However, the History and Philosophy of Astronomy course proves to be more challenging in terms of grading as opposed to other courses. Finally, our study reveals a promising direction for automating grading systems for MOOCs, especially in subjects with well-defined rubrics.
翻译:大规模开放在线课程(MOOC)为全球任何拥有计算机和互联网接入的人打开了免费教育的大门。尽管学习实现了民主化,但这些课程的大规模注册意味着一名讲师几乎不可能评估每位学生的写作作业。因此,通常依据简单评分标准的同行互评成为首选方法。虽然方便,但同行互评在可靠性和有效性方面往往存在不足。在本研究中,我们通过18种不同设定,探索利用大型语言模型(LLM)替代MOOC中同行互评的可行性。具体而言,我们聚焦两种最先进的LLM:GPT-4和GPT-3.5,并涉及三门不同课程:天文学导论、天体生物学、以及天文学历史与哲学。为指示LLM,我们基于零样本思维链(Zero-shot-CoT)提示技术的变体使用了三种不同提示:结合讲师提供正确答案的零样本思维链;同时结合讲师制定答案与评分标准的零样本思维链;以及结合讲师提供正确答案与LLM生成评分标准的零样本思维链。结果表明,当整合讲师提供的答案与评分标准时,零样本思维链产生的评分比同行互评与讲师评分更为一致。然而,天文学历史与哲学课程在评分方面相比其他课程更具挑战性。最后,本研究揭示了自动化MOOC评分系统的可行方向,尤其在具有明确评分标准的学科中。