Written answers to open-ended questions can have a higher long-term effect on learning than multiple-choice questions. However, it is critical that teachers immediately review the answers, and ask to redo those that are incoherent. This can be a difficult task and can be time-consuming for teachers. A possible solution is to automate the detection of incoherent answers. One option is to automate the review with Large Language Models (LLM). In this paper, we analyze the responses of fourth graders in mathematics using three LLMs: GPT-3, BLOOM, and YOU. We used them with zero, one, two, three and four shots. We compared their performance with the results of various classifiers trained with Machine Learning (ML). We found that LLMs perform worse than MLs in detecting incoherent answers. The difficulty seems to reside in recursive questions that contain both questions and answers, and in responses from students with typical fourth-grader misspellings. Upon closer examination, we have found that the ChatGPT model faces the same challenges.
翻译:开放式问题的书面答案对学习的长期效果可能优于选择题。然而,教师必须立即审阅这些答案,并要求重做那些不连贯的答案,这一任务既困难又耗时。一个可行的解决方案是自动检测不连贯的答案,而采用大语言模型(LLM)进行自动化审阅是一种途径。本文分析了四年级学生在数学题中的回答,使用了三种LLM:GPT-3、BLOOM和YOU,并分别通过零样本、一样本、两样本、三样本和四样本提示进行测试。我们将它们的性能与基于机器学习(ML)训练的各种分类器的结果进行了比较。实验发现,LLM在检测不连贯答案方面表现不如ML。困难似乎在于递归提问(同时包含问题和答案)以及含有典型四年级学生拼写错误的回答。进一步分析表明,ChatGPT模型也面临相同的挑战。