LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline

Generative AI and large language models (LLMs) are increasingly applied to question generation and automated assessment. However, deploying LLMs in preparation for high-stakes exams requires more than prompt engineering; it demands software pipelines that systematically ground model outputs in authorised curriculum artefacts and marking guidelines issued by education authorities. This paper presents a curriculum-grounded, configurable LLM-as-Judge pipeline for question-level marking, co-developed with an industrial partner, to support exam preparation for university admission. The pipeline identifies the relevant topics, subtopics, and cognitive demand of a question, and assembles verifiable and authorised context to support LLM judgement. Curriculum intent is operationalised through concrete syllabus artefacts, including prescribed verbs and outcomes, performance band descriptors, glossary definitions, and marking-guideline principles. A staged LLM workflow is employed to first generate question-specific rubrics, capturing structured expectations of performance, and then derive and evaluate marking criteria used to allocate marks to student responses. This design improves consistency, transparency, and alignment with official marking practices. Preliminary evaluation shows that the proposed LLM-as-Judge pipeline delivers marking outcomes comparable to human tutors, while yielding justifications that are more traceable to authorised curriculum artefacts and marking standards. The pipeline has also been integrated into an online study platform, where early deployment data provide initial insights into operational usage and manual overrides.

翻译：生成式人工智能与大语言模型（LLMs）正日益应用于试题生成与自动评估领域。然而，在高风险考试备考中部署LLMs，需要的不仅是提示工程，更需构建系统化的软件管线，使模型输出严格锚定教育主管部门授权的课程文档及评分标准。本文提出一种基于课程体系、可配置的LLM-as-Judge评分管线，用于试题级评分，该管线与产业合作伙伴协同开发，旨在支持大学入学考试备考。该管线能够识别试题涉及的主题、子主题及认知需求，并整合可验证且有授权的上下文信息以支撑LLM判断。课程意图通过具体化教学大纲文档（包括规定动词与学习成果、表现等级描述符、术语表定义及评分指导原则）加以实现。采用分阶段LLM工作流程：首先生成试题专属评分细则，捕捉结构化的预期表现；随后推导并评估用于学生作答评分的关键标准。此设计提升了评分一致性、透明度以及与官方评分实践的契合度。初步评估表明，所提出的LLM-as-Judge管线在产出与人类教师可比的评分结果的同时，能提供更可追溯至授权课程文档与评分标准的判分依据。该管线已集成至在线学习平台，早期部署数据初步揭示了运行使用与人工覆写的实证特征。

相关内容

课程

关注 6

课程是指学校学生所应学习的学科总和及其进程与安排。课程是对教育的目标、教学内容、教学活动方式的规划和设计，是教学计划、教学大纲等诸多方面实施过程的总和。广义的课程是指学校为实现培养目标而选择的教育内容及其进程的总和，它包括学校老师所教授的各门学科和有目的、有计划的教育活动。狭义的课程是指某一门学科。专知上对国内外最新AI+X的课程进行了收集与索引，涵盖斯坦福大学、CMU、MIT、清华、北大等名校开放课程。

LLM/智能体作为数据分析师：综述

专知会员服务

38+阅读 · 2025年9月30日

迈向LLM时代的可泛化评估：超越基准的综述

专知会员服务

23+阅读 · 2025年4月29日

大型语言模型（LLM）智能体全栈安全的综述：数据、训练与部署

专知会员服务

33+阅读 · 2025年4月23日

【新书】解码大型语言模型：理解、实现与优化LLM在自然语言处理应用中的全面指南

专知会员服务

49+阅读 · 2024年12月13日