As Competency-Based Education (CBE) is gaining traction around the world, the shift from marks-based assessment to qualitative competency mapping is a manual challenge for educators. This paper tackles the bottleneck issue by suggesting a "Human-in-the-Loop" benchmarking framework to assess the effectiveness of multiple LLMs in automating secondary-level mathematics assessment. Based on the Grade 10 Optional Mathematics curriculum in Nepal, we created a multi-dimensional rubric for four topics and four cross-cutting competencies: Comprehension, Knowledge, Operational Fluency, and Behavior and Correlation. The multi-provider ensemble, consisted of open-weight models -- Eagle (Llama 3.1-8B) and Orion (Llama 3.3-70B) -- and proprietary frontier models Nova (Gemini 2.5 Flash) and Lyra (Gemini 3 Pro), was benchmarked against a ground truth defined by two senior mathematics faculty members (kappa_w = 0.8652). The findings show a marked "Architecture-compatibility gap". Although the Gemini-based Mixture-of-Experts (Sparse MoE) models achieved "Fair Agreement" (kappa_w ~ 0.38), the larger Orion (70B) model exhibited "No Agreement" (kappa_w = -0.0261), suggesting that architectural compliance with instruction constraints outweighs the scale of raw parameters in rubric-constrained tasks. We conclude that while LLMs are not yet suitable for autonomous certification, they provide high-value assistive support for preliminary evidence extraction within a "Human-in-the-Loop" framework.
翻译:随着能力本位教育(CBE)在全球范围内逐渐普及,从基于分数的评估向定性能力图谱的转变给教育工作者带来了人工层面的挑战。本文通过提出一种"人在回路"基准测试框架来应对这一瓶颈问题,旨在评估多种大语言模型(LLMs)在自动化中学数学评估中的有效性。基于尼泊尔十年级数学选修课程,我们创建了一个涵盖四个主题和四项跨领域能力(理解力、知识掌握、运算流畅性、行为与相关性)的多维度评分标准。由开放权重模型Eagle(Llama 3.1-8B)与Orion(Llama 3.3-70B)以及专有前沿模型Nova(Gemini 2.5 Flash)与Lyra(Gemini 3 Pro)组成的多提供方集成模型,以两位高级数学教职人员定义的真实基准(kappa_w = 0.8652)进行对照评估。研究结果揭示了显著的"架构兼容性鸿沟":尽管基于Gemini的混合专家(稀疏MoE)模型达到了"公平一致性"(kappa_w ≈ 0.38),但规模更大的Orion(70B)模型却呈现"无一致性"(kappa_w = -0.0261),这表明在评分规则约束型任务中,架构对指令约束的适应性比原始参数规模更为关键。我们得出结论:虽然大语言模型尚不适用于自主认证,但在"人在回路"框架下,它们能为初步证据提取提供高价值的辅助支持。