Assessing teachers' geometric content knowledge is essential for geometry instructional quality and student learning, but difficult to scale. The Van Hiele model characterizes geometric reasoning through five hierarchical levels. Traditional Van Hiele assessment relies on manual expert analysis of open-ended responses. This process is time-consuming, costly, and prevents large-scale evaluation. This study develops an automated approach for diagnosing teachers' Van Hiele reasoning levels using large language models grounded in educational theory. Our central hypothesis is that integrating explicit skills information significantly improves Van Hiele classification. In collaboration with mathematics education researchers, we built a structured skills dictionary decomposing the Van Hiele levels into 33 fine-grained reasoning skills. Through a custom web platform, 31 pre-service teachers solved geometry problems, yielding 226 responses. Expert researchers then annotated each response with its Van Hiele level and demonstrated skills from the dictionary. Using this annotated dataset, we implemented two classification approaches: (1) retrieval-augmented generation (RAG) and (2) multi-task learning (MTL). Each approach compared a skills-aware variant incorporating the skills dictionary against a baseline without skills information. Results showed that for both methods, skills-aware variants significantly outperformed baselines across multiple evaluation metrics. This work provides the first automated approach for Van Hiele level classification from open-ended responses. It offers a scalable, theory-grounded method for assessing teachers' geometric reasoning that can enable large-scale evaluation and support adaptive, personalized teacher learning systems.
翻译:评估教师的几何内容知识对于几何教学质量和学生学习至关重要,但难以大规模开展。范希尔模型通过五个层级来表征几何推理能力。传统的范希尔评估依赖专家对开放性回答进行人工分析,这个过程耗时费力,且阻碍了大规模评估。本研究开发了一种基于教育理论的大语言模型自动诊断教师范希尔推理层级的方法。我们的核心假设是,整合明确的技能信息能显著改进范希尔层级分类。通过与数学教育研究者合作,我们构建了一个结构化的技能词典,将范希尔层级分解为33种细粒度推理技能。通过一个自定义的网页平台,31名职前教师解决了几何问题,共收集到226份回答。专家研究者随后对每份回答的范希尔层级及从词典中展示的技能进行了标注。利用这个标注数据集,我们实施了两种分类方法:(1) 检索增强生成 (RAG) 和 (2) 多任务学习 (MTL)。每种方法都将一个包含技能词典的“技能感知”变体与一个不包含技能信息的基线模型进行了比较。结果表明,对于两种方法,技能感知变体在多项评估指标上均显著优于基线模型。本研究首次提供了从开放性回答中自动进行范希尔层级分类的方法。它提供了一种可扩展、基于理论的评估教师几何推理能力的方法,能够支持大规模评估,并为自适应、个性化的教师学习系统提供支持。