In this paper, we explore the application of large language models (LLMs) for generating code-tracing questions in introductory programming courses. We designed targeted prompts for GPT4, guiding it to generate code-tracing questions based on code snippets and descriptions. We established a set of human evaluation metrics to assess the quality of questions produced by the model compared to those created by human experts. Our analysis provides insights into the capabilities and potential of LLMs in generating diverse code-tracing questions. Additionally, we present a unique dataset of human and LLM-generated tracing questions, serving as a valuable resource for both the education and NLP research communities. This work contributes to the ongoing dialogue on the potential uses of LLMs in educational settings.
翻译:本文探讨了大语言模型(LLM)在入门编程课程中生成代码追踪题目的应用。我们针对GPT4设计了特定的提示词,引导其基于代码片段和描述生成代码追踪题目。建立了一套人工评估指标,用于评估模型生成题目与人类专家编写题目的质量对比。我们的分析揭示了大语言模型在生成多样化代码追踪题目方面的能力与潜力。此外,我们提供了一个包含人类和LLM生成追踪题目的独特数据集,为教育界和自然语言处理研究界提供了宝贵资源。这项工作为关于大语言模型在教育场景中潜在应用的持续讨论做出了贡献。