Massive open online courses (MOOCs) unlock the doors to free education for anyone around the globe with access to a computer and the internet. Despite this democratization of learning, the massive enrollment in these courses means it is almost impossible for one instructor to assess every student's writing assignment. As a result, peer grading, often guided by a straightforward rubric, is the method of choice. While convenient, peer grading often falls short in terms of reliability and validity. In this study, using 18 distinct settings, we explore the feasibility of leveraging large language models (LLMs) to replace peer grading in MOOCs. Specifically, we focus on two state-of-the-art LLMs: GPT-4 and GPT-3.5, across three distinct courses: Introductory Astronomy, Astrobiology, and the History and Philosophy of Astronomy. To instruct LLMs, we use three different prompts based on a variant of the zero-shot chain-of-thought (Zero-shot-CoT) prompting technique: Zero-shot-CoT combined with instructor-provided correct answers; Zero-shot-CoT in conjunction with both instructor-formulated answers and rubrics; and Zero-shot-CoT with instructor-offered correct answers and LLM-generated rubrics. Our results show that Zero-shot-CoT, when integrated with instructor-provided answers and rubrics, produces grades that are more aligned with those assigned by instructors compared to peer grading. However, the History and Philosophy of Astronomy course proves to be more challenging in terms of grading as opposed to other courses. Finally, our study reveals a promising direction for automating grading systems for MOOCs, especially in subjects with well-defined rubrics.
翻译:大规模开放在线课程(MOOC)为全球任何拥有计算机和互联网接入的人们打开了免费教育的大门。尽管学习资源得以普及,但这些课程的庞大注册人数意味着,几乎不可能由一位教师评估每位学生的写作作业。因此,通常依据简单评分量规的同伴互评成为了首选方法。虽然便捷,但同伴互评在可靠性和有效性方面往往有所欠缺。在本研究中,我们利用18种不同设置,探索借助大型语言模型(LLM)取代MOOC中同伴互评的可行性。具体而言,我们聚焦于两种最先进的LLM:GPT-4和GPT-3.5,并将其应用于三门不同的课程:天文学导论、天体生物学以及天文学历史与哲学。为指导LLM,我们基于零样本思维链(Zero-shot-CoT)提示技术的变体,使用了三种不同提示:结合教师提供正确答案的零样本思维链;结合教师拟定答案与评分量规的零样本思维链;以及结合教师提供正确答案与LLM生成评分量规的零样本思维链。结果表明,与同伴互评相比,当零样本思维链与教师提供的答案及评分量规相结合时,其给出的分数与教师评定的分数更为一致。然而,天文学历史与哲学课程在评分方面相较其他课程更具挑战性。最后,本研究为MOOC评分系统的自动化(尤其在具有明确定义评分量规的学科中)揭示了有前景的发展方向。