This study evaluates the pedagogical viability of LLM-generated English as a Foreign Language (EFL) learning content. Utilising log data from Japanese junior high school students practicing on a grammar drilling application, we analysed how different question modalities impact student performance and whether theoretical localised CEFR difficulty tiers accurately predict empirical task difficulty. Results reveal a clear performance hierarchy: multiple-choice questions carried the lowest cognitive load, cloze tasks posed the greatest barrier to active recall, and drag-and-drop exercises incurred the heaviest time penalties. Furthermore, learner data validated the CEFR-J grammar framework, showing a steady decline in accuracy and increased response times as proficiency levels advanced. These findings demonstrate that LLMs can successfully generate learning content, while highlighting the need for developers to strategically sequence question modalities to transition learners from passive recognition to active linguistic production.
翻译:本研究评估了大语言模型(LLM)生成的英语作为外语(EFL)学习内容的教学可行性。通过利用日本初中生使用语法练习应用程序时的日志数据,我们分析了不同题目模态对学生表现的影响,以及理论上的本地化CEFR难度层级是否能准确预测实证任务难度。结果显示出一个明确的表现层级:选择题承载的认知负荷最低,填空任务对主动回忆构成最大障碍,而拖拽练习则带来最严重的时间代价。此外,学习者的数据验证了CEFR-J语法框架的有效性,表明随着语言水平的提升,准确率稳步下降,而响应时间逐渐增加。这些发现表明,LLM能够成功生成学习内容,同时强调了开发者需策略性地排列题目模态,以引导学习者从被动识别过渡到主动语言生成。