Assessing teachers' geometric content knowledge is essential for geometry instructional quality and student learning, but difficult to scale. The Van Hiele model characterizes geometric reasoning through five hierarchical levels. Traditional Van Hiele assessment relies on manual expert analysis of open-ended responses. This process is time-consuming, costly, and prevents large-scale evaluation. This study develops an automated approach for diagnosing teachers' Van Hiele reasoning levels using large language models grounded in educational theory. Our central hypothesis is that integrating explicit skills information significantly improves Van Hiele classification. In collaboration with mathematics education researchers, we built a structured skills dictionary decomposing the Van Hiele levels into 33 fine-grained reasoning skills. Through a custom web platform, 31 pre-service teachers solved geometry problems, yielding 226 responses. Expert researchers then annotated each response with its Van Hiele level and demonstrated skills from the dictionary. Using this annotated dataset, we implemented two classification approaches: (1) retrieval-augmented generation (RAG) and (2) multi-task learning (MTL). Each approach compared a skills-aware variant incorporating the skills dictionary against a baseline without skills information. Results showed that for both methods, skills-aware variants significantly outperformed baselines across multiple evaluation metrics. This work provides the first automated approach for Van Hiele level classification from open-ended responses. It offers a scalable, theory-grounded method for assessing teachers' geometric reasoning that can enable large-scale evaluation and support adaptive, personalized teacher learning systems.


翻译:评估教师的几何内容知识对于几何教学质量和学生学习至关重要,但难以大规模开展。范希尔模型通过五个层级来表征几何推理能力。传统的范希尔评估依赖专家对开放性回答进行人工分析,这个过程耗时费力,且阻碍了大规模评估。本研究开发了一种基于教育理论的大语言模型自动诊断教师范希尔推理层级的方法。我们的核心假设是,整合明确的技能信息能显著改进范希尔层级分类。通过与数学教育研究者合作,我们构建了一个结构化的技能词典,将范希尔层级分解为33种细粒度推理技能。通过一个自定义的网页平台,31名职前教师解决了几何问题,共收集到226份回答。专家研究者随后对每份回答的范希尔层级及从词典中展示的技能进行了标注。利用这个标注数据集,我们实施了两种分类方法:(1) 检索增强生成 (RAG) 和 (2) 多任务学习 (MTL)。每种方法都将一个包含技能词典的“技能感知”变体与一个不包含技能信息的基线模型进行了比较。结果表明,对于两种方法,技能感知变体在多项评估指标上均显著优于基线模型。本研究首次提供了从开放性回答中自动进行范希尔层级分类的方法。它提供了一种可扩展、基于理论的评估教师几何推理能力的方法,能够支持大规模评估,并为自适应、个性化的教师学习系统提供支持。

0
下载
关闭预览

相关内容

大模型数学推理数据合成相关方法
专知会员服务
36+阅读 · 2025年1月19日
自动结构变分推理,Automatic structured variational inference
专知会员服务
41+阅读 · 2020年2月10日
【课程推荐】 深度学习中的几何(Geometry of Deep Learning)
专知会员服务
59+阅读 · 2019年11月10日
知识图谱的自动构建
DataFunTalk
58+阅读 · 2019年12月9日
人工智能在教育领域的应用探析
MOOC
14+阅读 · 2019年3月16日
基于面部表情的学习困惑自动识别法
MOOC
10+阅读 · 2018年9月17日
论文浅尝 | 基于神经网络的知识推理
开放知识图谱
15+阅读 · 2018年3月12日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
25+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
26+阅读 · 2011年12月31日
Arxiv
0+阅读 · 3月27日
VIP会员
最新内容
《无人机对海面作战影响评估》
专知会员服务
1+阅读 · 今天15:30
印度精确打击与指挥架构的断层
专知会员服务
4+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
相关VIP内容
大模型数学推理数据合成相关方法
专知会员服务
36+阅读 · 2025年1月19日
自动结构变分推理,Automatic structured variational inference
专知会员服务
41+阅读 · 2020年2月10日
【课程推荐】 深度学习中的几何(Geometry of Deep Learning)
专知会员服务
59+阅读 · 2019年11月10日
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
25+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
26+阅读 · 2011年12月31日
Top
微信扫码咨询专知VIP会员