Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessment questions that are aligned with pedagogical frameworks (e.g., Bloom's taxonomy). However, they often rely on subjective or limited evaluation methods; focus primarily on proprietary models; or rarely systematically examine generation, evaluation, or deployment constraints in real educational settings. Meanwhile, Small Language Models (SLMs) have emerged as local alternatives that better address privacy and resource limitations; yet their effectiveness for assessment tasks remains underexplored. To address this gap, we systematically compare LLMs and SLMs for assessment question design; evaluate generation quality across Bloom's taxonomy levels using reproducible, pedagogically grounded metrics; and further assess model-based judging against expert-informed evaluation by analyzing reliability and agreement patterns. Results show that SLMs achieve competitive performance across key pedagogically motivated quality dimensions while enabling local, privacy-sensitive deployment. However, model-based evaluations also exhibit systematic inconsistencies and bias relative to expert ratings. These findings provide evidence to posit language models as bounded assistants in assessment workflows; underscore the necessity of Human-in-the-Loop; and advance the automated educational question generation field by examining quality, reliability, and deployment-aware trade-offs.
翻译:生成式人工智能日益支持教育设计任务,例如通过大型语言模型(LLMs),展示了设计符合教学框架(如布鲁姆分类法)的评估问题的能力。然而,它们往往依赖主观或有限评估方法;主要关注专有模型;或很少系统性地检查真实教育环境中的生成、评估或部署约束。同时,小规模语言模型(SLMs)作为本地化替代方案出现,能更好地解决隐私和资源限制问题;然而其在评估任务中的有效性仍未充分探索。为填补这一空白,我们系统性地比较了LLMs和SLMs在评估问题设计上的表现;使用可复现、基于教学法指标的度量标准,评估了跨布鲁姆分类法层级的生成质量;并通过分析一致性和一致性模式,进一步评估了基于模型的评判与专家评审的对比。结果表明,SLMs在关键教学驱动的质量维度上达到了有竞争力的性能,同时支持本地化、保护隐私的部署。然而,基于模型的评估相对于专家评分也表现出系统性不一致性和偏差。这些发现为将语言模型定位为评估工作流中的有限助手提供了证据;强调了人在回路中的必要性;并通过考察质量、可靠性和部署相关权衡,推动了自动化教育问题生成领域的发展。