This paper introduces the first publicly available dataset for Automatic Essay Scoring (AES) and feedback generation in Basque, targeting the CEFR C1 proficiency level. The dataset comprises 3,200 essays from HABE, each annotated by expert evaluators with criterion specific scores covering correctness, richness, coherence, cohesion, and task alignment enriched with detailed feedback and error examples. We fine-tune open-source models, including RoBERTa-EusCrawl and Latxa 8B/70B, for both scoring and explanation generation. Our experiments show that encoder models remain highly reliable for AES, while supervised fine-tuning (SFT) of Latxa significantly enhances performance, surpassing state-of-the-art (SoTA) closed-source systems such as GPT-5 and Claude Sonnet 4.5 in scoring consistency and feedback quality. We also propose a novel evaluation methodology for assessing feedback generation, combining automatic consistency metrics with expert-based validation of extracted learner errors. Results demonstrate that the fine-tuned Latxa model produces criterion-aligned, pedagogically meaningful feedback and identifies a wider range of error types than proprietary models. This resource and benchmark establish a foundation for transparent, reproducible, and educationally grounded NLP research in low-resource languages such as Basque.
翻译:本文介绍了首个面向巴斯克语公开可用的自动作文评分(AES)及反馈生成数据集,针对欧洲共同语言参考标准(CEFR)C1水平。该数据集包含来自HABE的3200篇作文,每篇均由专家评估员按照特定标准进行评分,涵盖正确性、丰富度、连贯性、衔接性及任务契合度,并配有详细反馈和错误示例。我们针对评分与解释生成两种任务,微调了包括RoBERTa-EusCrawl与Latxa 8B/70B在内的开源模型。实验表明,编码器模型在AES任务上仍保持高可靠性,而Latxa模型的监督微调(SFT)显著提升性能,在评分一致性与反馈质量上超越GPT-5、Claude Sonnet 4.5等当前最优闭源系统(SoTA)。我们还提出一种评估反馈生成的新型方法论,结合自动一致性指标与基于专家验证的学习者错误提取结果。结果显示,微调后的Latxa模型能够生成符合评分标准、具有教学意义的反馈,并识别出比专有模型更广泛的错误类型。该资源与基准为巴斯克语等低资源语言中透明、可复现及教育导向的自然语言处理研究奠定了基础。