(Tack et al., 2023) organized the shared task hosted by the 18th Workshop on Innovative Use of NLP for Building Educational Applications on generation of teacher language in educational dialogues. Following the structure of the shared task, in this study, we attempt to assess the generative abilities of large language models in providing informative and helpful insights to students, thereby simulating the role of a knowledgeable teacher. To this end, we present an extensive evaluation of several benchmarking generative models, including GPT-4 (few-shot, in-context learning), fine-tuned GPT-2, and fine-tuned DialoGPT. Additionally, to optimize for pedagogical quality, we fine-tuned the Flan-T5 model using reinforcement learning. Our experimental findings on the Teacher-Student Chatroom Corpus subset indicate the efficacy of GPT-4 over other fine-tuned models, measured using BERTScore and DialogRPT. We hypothesize that several dataset characteristics, including sampling, representativeness, and dialog completeness, pose significant challenges to fine-tuning, thus contributing to the poor generalizability of the fine-tuned models. Finally, we note the need for these generative models to be evaluated with a metric that relies not only on dialog coherence and matched language modeling distribution but also on the model's ability to showcase pedagogical skills.
翻译:(Tack等人,2023)组织了由第18届自然语言处理在教育应用创新研讨会主办的共享任务,旨在教育对话中生成教师语言。遵循该共享任务的结构,本研究尝试评估大型语言模型在向学生提供信息丰富且有用的见解方面的生成能力,从而模拟知识渊博教师的角色。为此,我们对多个基准生成模型进行了广泛评估,包括GPT-4(少样本、上下文学习)、微调GPT-2以及微调DialoGPT。此外,为优化教学质量,我们使用强化学习对Flan-T5模型进行了微调。在师生聊天语料库子集上的实验结果表明,GPT-4相比其他微调模型更有效,通过BERTScore和DialogRPT进行评估。我们假设,包括采样、代表性和对话完整性在内的多个数据集特征对微调构成了重大挑战,从而导致了微调模型的泛化能力较差。最后,我们指出,这些生成模型需采用不仅依赖对话连贯性和匹配语言建模分布,还应体现模型展示教学技能的评估指标。