As Competency-Based Education (CBE) is gaining traction around the world, the shift from marks-based assessment to qualitative competency mapping is a manual challenge for educators. This paper tackles the bottleneck issue by suggesting a "Human-in-the-Loop" benchmarking framework to assess the effectiveness of multiple LLMs in automating secondary-level mathematics assessment. Based on the Grade 10 Optional Mathematics curriculum in Nepal, we created a multi-dimensional rubric for four topics and four cross-cutting competencies: Comprehension, Knowledge, Operational Fluency, and Behavior and Correlation. The multi-provider ensemble, consisted of open-weight models -- Eagle (Llama 3.1-8B) and Orion (Llama 3.3-70B) -- and proprietary frontier models Nova (Gemini 2.5 Flash) and Lyra (Gemini 3 Pro), was benchmarked against a ground truth defined by two senior mathematics faculty members (kappa_w = 0.8652). The findings show a marked "Architecture-compatibility gap". Although the Gemini-based Mixture-of-Experts (Sparse MoE) models achieved "Fair Agreement" (kappa_w ~ 0.38), the larger Orion (70B) model exhibited "No Agreement" (kappa_w = -0.0261), suggesting that architectural compliance with instruction constraints outweighs the scale of raw parameters in rubric-constrained tasks. We conclude that while LLMs are not yet suitable for autonomous certification, they provide high-value assistive support for preliminary evidence extraction within a "Human-in-the-Loop" framework.


翻译:随着能力本位教育(CBE)在全球范围内逐渐普及,从基于分数的评估向定性能力图谱的转变给教育工作者带来了人工层面的挑战。本文通过提出一种"人在回路"基准测试框架来应对这一瓶颈问题,旨在评估多种大语言模型(LLMs)在自动化中学数学评估中的有效性。基于尼泊尔十年级数学选修课程,我们创建了一个涵盖四个主题和四项跨领域能力(理解力、知识掌握、运算流畅性、行为与相关性)的多维度评分标准。由开放权重模型Eagle(Llama 3.1-8B)与Orion(Llama 3.3-70B)以及专有前沿模型Nova(Gemini 2.5 Flash)与Lyra(Gemini 3 Pro)组成的多提供方集成模型,以两位高级数学教职人员定义的真实基准(kappa_w = 0.8652)进行对照评估。研究结果揭示了显著的"架构兼容性鸿沟":尽管基于Gemini的混合专家(稀疏MoE)模型达到了"公平一致性"(kappa_w ≈ 0.38),但规模更大的Orion(70B)模型却呈现"无一致性"(kappa_w = -0.0261),这表明在评分规则约束型任务中,架构对指令约束的适应性比原始参数规模更为关键。我们得出结论:虽然大语言模型尚不适用于自主认证,但在"人在回路"框架下,它们能为初步证据提取提供高价值的辅助支持。

0
下载
关闭预览

相关内容

数学是关于数量、结构、变化等主题的探索。
大语言模型智能体的评估与基准:综述
专知会员服务
50+阅读 · 2025年7月31日
【斯坦福博士论文】大语言模型的AI辅助评估
专知会员服务
31+阅读 · 2025年3月30日
《以人为中心的大型语言模型(LLM)研究综述》
专知会员服务
41+阅读 · 2024年11月25日
天大最新《大型语言模型评估》全面综述,111页pdf
专知会员服务
89+阅读 · 2023年10月31日
推荐|上交大推出Texygen:文本生成模型的基准测试平台
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
1+阅读 · 2016年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
28+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
VIP会员
最新内容
《无人机对海面作战影响评估》
专知会员服务
1+阅读 · 43分钟前
印度精确打击与指挥架构的断层
专知会员服务
4+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
相关基金
国家自然科学基金
1+阅读 · 2016年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
28+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
Top
微信扫码咨询专知VIP会员